If the platform forbids agents, build the dataset yourself from the source
Filings, financial statements and earnings calls were public before any vendor packaged them. An agent can go back to the source, and the tools to do it are free.
A license can stop you from putting a vendor’s files into an AI system. It cannot stop you from collecting the same public facts yourself. Until now that was too much work for one researcher. With an agent it is a week.
What you can build from the source
| Dataset | Public source | What the agent does | What you do not get |
|---|---|---|---|
| Financial statements | SEC XBRL APIs: every reported concept for a company in one call | Builds a firm-quarter panel and documents each field | History from before XBRL reporting, and a vendor’s standardized items |
| Filing text | SEC submissions history and the filings themselves | Downloads, splits into sections and cleans 10-K, 10-Q and 8-K text | Firms that do not file with the SEC |
| Call transcripts | Company webcasts and replays | Transcribes with open speech recognition, splits speakers, tags the Q&A | The past. Replays expire, and S&P’s history goes back to 2004 |
| Macro and banking series | FRED, the BIS, the ECB, the World Bank | Pulls and documents each series | Little. These are already free |
| Factors and anomalies | Open Source Asset Pricing, the French library, Global Factor Data | Downloads and merges | The security-level returns underneath |
| Security prices and returns | Nothing open matches CRSP | This is the one you still license |
The SEC describes its data service plainly: the APIs “do not require any authentication or API keys to access,” they cover filing history and the XBRL data from financial statements, and the files “are updated throughout the day, in real time.” A bulk download of everything is republished every night.
The transcript example
S&P lists more than 24,000 entities in its transcript product with history back to 2004. You cannot rebuild that. What you can build is the part most new projects need: recent calls, for the firms in your sample.
- Collect the webcast or replay from each company’s investor relations page.
- Transcribe it with an open model. Whisper’s code and weights are under the MIT License and it handles many languages, which matters for firms outside the US.
- Structure the text: speakers, prepared remarks, questions and answers.
- Check a sample by ear and report the error rate in the paper.
- Keep the audio references and a manifest, so anyone can rebuild the set.
The result is yours. No clause governs what an agent may do with it, and you can build measures on it with any model you like.
Open source tools that already do this
These projects wrap public data sources so an agent can query them. Star counts are from GitHub on 9 October 2026.
| Project | What it reaches | License | Stars |
|---|---|---|---|
stefanoamorelli/sec-edgar-mcp |
SEC EDGAR filings | AGPL-3.0 | 368 |
daniel3303/Equibles |
Self-hosted SEC filings and XBRL financials | AGPL-3.0 | 230 |
stefanoamorelli/fred-mcp-server |
FRED economic data | AGPL-3.0 | 124 |
lzinga/us-gov-open-data-mcp |
More than 40 US government data APIs | MIT | 112 |
hanlulong/openecon-data |
FRED, World Bank, IMF and Eurostat indicators | See repository | 84 |
FTShare-Lab/FTShare-MCP |
Chinese market data and factors | MIT | 250 |
Read our warning on other people’s MCP servers before you install any of them. Each one wraps an API you can also call directly, and for a dataset that will sit under a paper, direct is better: fewer moving parts and nothing between you and the source.
The source has rules too
Going to the source does not mean no rules. The SEC limits each user to 10 requests a second and says it does not allow “unclassified” bots to crawl its site, so identify your requests and stay under the limit. Company webcasts carry their own terms. And building your own transcript is different from redistributing a vendor’s.
What it costs you
You become the vendor. Coverage gaps, transcription errors, delisted firms and identifier mistakes are now your problem, and a referee will ask about each one. The hard part is not the download. It is linking: the SEC’s company key has to be matched to the identifiers in your returns data, and nobody gives that away.
So build what is public, license what is scarce, and say in the paper which is which. A department that can build most of a dataset negotiates very differently for the rest.
Sources
- SEC sec.gov
- SEC, developer resources and access limits sec.gov
- Whisper, open source speech recognition github.com
- S&P Global Marketplace, Machine Readable Transcripts marketplace.spglobal.com
- Open Source Asset Pricing openassetpricing.com