Analysis

Should you still pay for call transcripts when an agent can do the parsing?

An agent now does the cleaning, splitting and scoring that used to justify a data budget. What it cannot produce is history, identifiers and the right to use the text.

AI Fin ResearchCovers Global

The question in the headline is a claim to test. Our answer: yes for transcripts, for now, and for different reasons than before. The same test applies to any dataset you buy.

24,000+entities under coverage in S&P's Machine Readable Transcripts dataset, by its own listing
2004how far back that listing says the history goes
0years of that history an agent can recreate for you

What you were paying for

S&P’s listing describes transcripts of earnings, M&A, guidance, shareholder, conference and special calls, with speaker identifiers that link to its estimates data, and entity recognition that tags which companies are mentioned and where. The announcement of the same data on WRDS cites 9,400+ companies, coverage from 2000 in North America and 2004 globally, and tagging by company, speaker and key development.

Split that into parts and ask of each one whether an agent replaces it.

What the product gives you Does an agent replace it? Why
Splitting a call into presentation and Q&A, speakers and turns Yes An afternoon of work, and you can check it
Sentiment, uncertainty, topic and tone measures Yes And you should build them yourself, on more than one model, because the numbers change with the model
New calls, as they happen Partly Many calls are webcast, and speech recognition is good. Coverage and accuracy are then your problem
Twenty years of history No The calls happened once. Nobody can re-record 2009
Speaker identifiers linked to analysts and executives No Slow, manual and full of name collisions
Company identifiers that merge with returns and accounting data No This is what makes a text measure usable in a regression
The right to use the text in research No A license, not a file

The first two rows are what vendors used to be admired for. The last four are what they sell.

The clause that decides it

None of this helps if your license does not let an agent read the text. University libraries are explicit. The University of Waterloo says its agreements with publishers “do not allow sharing licensed materials with third parties,” including generative AI services. American University says its contracts “explicitly forbid uploading, processing, or otherwise using the content in AI systems,” and that this applies to all AI tools, including the ones the university itself licenses.

So a researcher can hold a transcript license and still not be allowed to run the measure every recent paper runs. Check before you build.

Buy or build

Your project Do this
Panel regressions over many years, merged with returns Buy. You need the history and the identifiers
A recent event window for a few hundred firms Building from public webcasts is feasible. Document the error rate
A new text measure on existing licensed transcripts Buy the data, build the measure yourself, and get the AI clause in writing
A dashboard or summary of what the calls said Do not buy anything for this. An agent does it

The same test for any data

Pay for what is slow to build and legally scarce: history, identifiers, linking tables, rights. Stop paying for what an agent does in an afternoon: cleaning, reshaping, summarizing, charting. When a renewal quote arrives, ask the vendor which of the two you are being charged for. If the answer is the interface, you have your negotiating position.

Sources

Related