Analysis

If the platform forbids agents, build the dataset yourself from the source

Filings, financial statements and earnings calls were public before any vendor packaged them. An agent can go back to the source, and the tools to do it are free.

AI Fin ResearchCovers US
If the platform says no agents, go around it to the source. The filings, the calls and the statistics were public before anyone packaged them.
0API keys needed for the SEC's filing and financial statement data
10requests a second the SEC allows each user
MITthe license on Whisper, the open speech recognition model that turns call audio into text

A license can stop you from putting a vendor’s files into an AI system. It cannot stop you from collecting the same public facts yourself. Until now that was too much work for one researcher. With an agent it is a week.

What you can build from the source

Dataset Public source What the agent does What you do not get
Financial statements SEC XBRL APIs: every reported concept for a company in one call Builds a firm-quarter panel and documents each field History from before XBRL reporting, and a vendor’s standardized items
Filing text SEC submissions history and the filings themselves Downloads, splits into sections and cleans 10-K, 10-Q and 8-K text Firms that do not file with the SEC
Call transcripts Company webcasts and replays Transcribes with open speech recognition, splits speakers, tags the Q&A The past. Replays expire, and S&P’s history goes back to 2004
Macro and banking series FRED, the BIS, the ECB, the World Bank Pulls and documents each series Little. These are already free
Factors and anomalies Open Source Asset Pricing, the French library, Global Factor Data Downloads and merges The security-level returns underneath
Security prices and returns Nothing open matches CRSP This is the one you still license

The SEC describes its data service plainly: the APIs “do not require any authentication or API keys to access,” they cover filing history and the XBRL data from financial statements, and the files “are updated throughout the day, in real time.” A bulk download of everything is republished every night.

The transcript example

S&P lists more than 24,000 entities in its transcript product with history back to 2004. You cannot rebuild that. What you can build is the part most new projects need: recent calls, for the firms in your sample.

  1. Collect the webcast or replay from each company’s investor relations page.
  2. Transcribe it with an open model. Whisper’s code and weights are under the MIT License and it handles many languages, which matters for firms outside the US.
  3. Structure the text: speakers, prepared remarks, questions and answers.
  4. Check a sample by ear and report the error rate in the paper.
  5. Keep the audio references and a manifest, so anyone can rebuild the set.

The result is yours. No clause governs what an agent may do with it, and you can build measures on it with any model you like.

Open source tools that already do this

These projects wrap public data sources so an agent can query them. Star counts are from GitHub on 9 October 2026.

Project What it reaches License Stars
stefanoamorelli/sec-edgar-mcp SEC EDGAR filings AGPL-3.0 368
daniel3303/Equibles Self-hosted SEC filings and XBRL financials AGPL-3.0 230
stefanoamorelli/fred-mcp-server FRED economic data AGPL-3.0 124
lzinga/us-gov-open-data-mcp More than 40 US government data APIs MIT 112
hanlulong/openecon-data FRED, World Bank, IMF and Eurostat indicators See repository 84
FTShare-Lab/FTShare-MCP Chinese market data and factors MIT 250

Read our warning on other people’s MCP servers before you install any of them. Each one wraps an API you can also call directly, and for a dataset that will sit under a paper, direct is better: fewer moving parts and nothing between you and the source.

The source has rules too

Going to the source does not mean no rules. The SEC limits each user to 10 requests a second and says it does not allow “unclassified” bots to crawl its site, so identify your requests and stay under the limit. Company webcasts carry their own terms. And building your own transcript is different from redistributing a vendor’s.

What it costs you

You become the vendor. Coverage gaps, transcription errors, delisted firms and identifier mistakes are now your problem, and a referee will ask about each one. The hard part is not the download. It is linking: the SEC’s company key has to be matched to the identifiers in your returns data, and nobody gives that away.

So build what is public, license what is scarce, and say in the paper which is which. A department that can build most of a dataset negotiates very differently for the rest.

Sources

Related