Should you still pay for call transcripts when an agent can do the parsing?
An agent now does the cleaning, splitting and scoring that used to justify a data budget. What it cannot produce is history, identifiers and the right to use the text.
The question in the headline is a claim to test. Our answer: yes for transcripts, for now, and for different reasons than before. The same test applies to any dataset you buy.
What you were paying for
S&P’s listing describes transcripts of earnings, M&A, guidance, shareholder, conference and special calls, with speaker identifiers that link to its estimates data, and entity recognition that tags which companies are mentioned and where. The announcement of the same data on WRDS cites 9,400+ companies, coverage from 2000 in North America and 2004 globally, and tagging by company, speaker and key development.
Split that into parts and ask of each one whether an agent replaces it.
| What the product gives you | Does an agent replace it? | Why |
|---|---|---|
| Splitting a call into presentation and Q&A, speakers and turns | Yes | An afternoon of work, and you can check it |
| Sentiment, uncertainty, topic and tone measures | Yes | And you should build them yourself, on more than one model, because the numbers change with the model |
| New calls, as they happen | Partly | Many calls are webcast, and speech recognition is good. Coverage and accuracy are then your problem |
| Twenty years of history | No | The calls happened once. Nobody can re-record 2009 |
| Speaker identifiers linked to analysts and executives | No | Slow, manual and full of name collisions |
| Company identifiers that merge with returns and accounting data | No | This is what makes a text measure usable in a regression |
| The right to use the text in research | No | A license, not a file |
The first two rows are what vendors used to be admired for. The last four are what they sell.
The clause that decides it
None of this helps if your license does not let an agent read the text. University libraries are explicit. The University of Waterloo says its agreements with publishers “do not allow sharing licensed materials with third parties,” including generative AI services. American University says its contracts “explicitly forbid uploading, processing, or otherwise using the content in AI systems,” and that this applies to all AI tools, including the ones the university itself licenses.
So a researcher can hold a transcript license and still not be allowed to run the measure every recent paper runs. Check before you build.
Buy or build
| Your project | Do this |
|---|---|
| Panel regressions over many years, merged with returns | Buy. You need the history and the identifiers |
| A recent event window for a few hundred firms | Building from public webcasts is feasible. Document the error rate |
| A new text measure on existing licensed transcripts | Buy the data, build the measure yourself, and get the AI clause in writing |
| A dashboard or summary of what the calls said | Do not buy anything for this. An agent does it |
The same test for any data
Pay for what is slow to build and legally scarce: history, identifiers, linking tables, rights. Stop paying for what an agent does in an afternoon: cleaning, reshaping, summarizing, charting. When a renewal quote arrives, ask the vendor which of the two you are being charged for. If the answer is the interface, you have your negotiating position.
Sources
- S&P Global Marketplace marketplace.spglobal.com
- WRDS data announcement wrds-www.wharton.upenn.edu
- University of Waterloo Libraries, use of library resources with AI uwaterloo.ca
- American University Library, uploading resources into AI tools answers.library.american.edu