The measurement frontier in climate finance has moved to text. The most influential firm-level climate exposure measures are now built by reading what companies say: machine-learning methods that scan earnings call transcripts for climate-related language, and transformer models fine-tuned to classify climate disclosures. These tools have reshaped how researchers quantify transition risk, and they have done so almost entirely on English-language corpora from developed markets.
I work in a market where the corpus is different. Not slightly different, but structurally different, in ways that determine what these methods can and cannot see. This note is an attempt to be precise about that difference, because I think it matters more than the literature currently acknowledges, and because the response it calls for is not the obvious one.
What the method assumes
The transcript-based approach rests on a specific institutional object: the quarterly earnings conference call. It assumes the call happens, that it is transcribed and commercially distributed, that it is conducted in English, and, this is the subtle part, that it contains a question-and-answer session. The Q&A matters because it is the least scriptable portion of corporate speech. Management can polish a prepared statement; it cannot fully rehearse an analyst's follow-up. Spontaneity is what makes the text informative about attention rather than presentation.
Every one of those assumptions is a feature of developed-market disclosure infrastructure. None of them is a law of nature.
What Jakarta produces instead
The Indonesian exchange lists more than nine hundred companies. A subset of them, concentrated in the large banks, telecoms, and consumer names that dominate the index, hold regular analyst calls, often in English, because their shareholder registers are international. Below that tier, the picture changes quickly. Many issuers hold no analyst calls at all. The disclosure event the exchange actually mandates is the paparan publik, an annual public expose whose materials are filed to the exchange and conducted predominantly in Bahasa Indonesia. Commercial transcripts of these events are, in my experience, rare to nonexistent.
And in the calls that do happen, there is the phenomenon anyone who has sat through one will recognise: code-switching. A CFO answers an analyst's English question and pivots mid-sentence into Bahasa Indonesia for the operational detail, or the reverse. The technical vocabulary of Indonesia's energy transition is itself hybrid: transisi energi, hilirisasi, PLTU for coal-fired plants, the biodiesel mandates known by codes like B35. A bigram dictionary built from S&P 500 calls does not contain these terms, and a translated dictionary misses how they are actually used.
| Corpus the method was built on | Corpus this market produces | |
|---|---|---|
| Event | Quarterly earnings call | Annual public expose; calls for large caps only |
| Language | English | Bahasa Indonesia, English, and code-switched mixtures |
| Transcription | Commercial, systematic | Rare; filed slide decks and minutes instead |
| Register | Prepared remarks plus spontaneous Q&A | Predominantly prepared material |
| Coverage | Near-universal for listed firms | Skewed toward the largest issuers |
Three failure modes
The first is selection. If exposure can only be measured where English transcripts exist, then the measurable universe is the large-cap tier, the firms with international investor bases and the resources to talk fluently about climate. The firms that fall out of the sample are the mid-caps in coal logistics, plantation supply chains, and energy services, which is to say a substantial share of where Indonesian transition risk actually sits. The measurement error is not random. It is concentrated precisely where the risk is.
The second is language. Keyword and bigram methods are brittle to code-switching, and brittle in a specific direction: they undercount. A firm discussing its coal phase-out in Bahasa Indonesia registers as a firm not discussing climate at all. The measured exposure of an Indonesian issuer is partly a function of which language its management happened to answer in.
The third is register. One could respond to the first two problems by switching corpus, moving from calls to the documents that exist market-wide: annual reports, sustainability reports mandated under POJK 51, public expose filings. Coverage improves dramatically. But these are prepared texts, and the literature on climate disclosure has a name for what prepared texts tend to contain. The evidence from transformer-based classification of corporate climate reporting is that voluntary, polished disclosure skews heavily toward non-committal language, or cheap talk. Moving from spontaneous speech to prepared documents trades a selection problem for a sincerity problem. Neither is obviously the smaller error.
What text can and cannot be a measure of
Working through this has changed how I frame the problem. The goal is usually stated as building a better text measure. But whatever a text-based method captures, what it captures directly is a firm's reporting practice. That is a real object and a useful one. It is not the same object as the firm's physical exposure, and the two can move apart: disclosure quality can improve substantially while emissions and asset-level hazard do not move at all.
This has a consequence I did not appreciate at first. A measure built from text needs a validation target that sits outside text. Without one, a firm that learns to write well about climate becomes, by construction, a firm with high measured climate attention, and there is no way to tell that outcome apart from a firm that has actually changed. The same trap appears whenever a measure is validated against another constructed measure rather than against something realised.
The honest counterargument
There is a reasonable objection: perhaps the large caps are where measurement matters most. They carry the index weight, the institutional ownership, the ESG-mandated capital. If the method covers them, it covers what is priced.
I do not think this survives contact with how transition risk propagates here. The listed banks' climate exposure is substantially their borrowers' exposure, and those borrowers include exactly the mid-tier firms the transcript universe misses. Measuring the bank while ignoring the borrower is measuring the mirror instead of the object. For an economy where coal, palm oil, and nickel processing run through dense supply chains of listed mid-caps, the undersampled tier is not noise around the signal. It is the signal.
What would actually work
The building blocks for the language problem exist and have simply not been assembled for this purpose. Indonesian NLP has mature pretrained models and benchmarks, the IndoBERT family and the IndoNLU resources, developed by a research community that took the language's specifics seriously. The mandatory disclosure corpus is public, complete across issuers, and growing richer under POJK 51. What does not yet exist, as far as I am aware, is a climate-finance lexicon and classifier built natively for Bahasa Indonesia and for code-switched financial speech, built from Indonesian regulatory filings and expose documents rather than translated from English seed terms.
That is a bounded, feasible piece of infrastructure. But I no longer think a better lexicon is sufficient on its own, and the reason follows from the selection problem above. If the firms that fall out of the text corpus are the mid-tier emitters, no amount of linguistic sophistication recovers them, because the missing input is not language. It is speech itself. Those firms are not talking.
What does not go silent is the physical world. Facilities are observable from orbit whether or not management holds a call. Asset-level emissions inventories built from satellite observation, facility coordinates intersected with physical hazard maps, and supply chain carbon intensity derived from industry-level input-output tables all have full cross-sectional coverage by construction, because none of them requires the firm to participate. Text then becomes one channel among several rather than the measurement itself, and it can be used for what it is uniquely good at: capturing what management attends to, and how that attention diverges from what the physical record shows.
Two design criteria follow, and I hold to both in my own work. The first is independence: no input to a climate exposure measure should be sourced from an ESG rating provider, because a measure that borrows from a rating cannot later be used to evaluate one. The second is external validation: the measure must be tested against realised outcomes rather than against other scores, since agreement between two constructed measures establishes nothing about either.
The tools exist. The corpus exists. The satellite record exists. What has not been built is the bridge between them for this market. That is not a complaint about the literature. It is a research agenda, and it is the one I am pursuing.
Related Reading
Sautner, Z., van Lent, L., Vilkov, G., and Zhang, R. (2023). Firm-Level Climate Change Exposure. Journal of Finance, 78(3), 1449–1498.
Bingler, J. A., Kraus, M., Leippold, M., and Webersinke, N. (2022). Cheap Talk and Cherry-Picking: What ClimateBert Has to Say on Corporate Climate Risk Disclosures. Finance Research Letters, 47.
Wilie, B., Vincentio, K., Winata, G. I., et al. (2020). IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding. Proceedings of AACL-IJCNLP.