Guide

Automate SEC filings research with a coding agent, from EDGAR to a finished panel

The SEC's data service needs no key. Here are the endpoints, the instruction file, a skill and a scheduled run, plus linking to CRSP and Compustat through the wrds library.

AI Fin ResearchWorking5 min read
Four endpoints, two small files and one command. That is the whole distance between you and an automated filings pipeline.

The endpoints

The SEC says its data APIs “do not require any authentication or API keys to access” and are “updated throughout the day, in real time.”

What you want Endpoint What comes back
Ticker to company key (CIK) https://www.sec.gov/files/company_tickers.json Every ticker with its CIK and name
A company’s filing history https://data.sec.gov/submissions/CIK##########.json Form types, dates, accession numbers, document names
Everything a company reported in XBRL https://data.sec.gov/api/xbrl/companyfacts/CIK##########.json Every concept, every period, in one call
One concept across all companies https://data.sec.gov/api/xbrl/frames/us-gaap/Assets/USD/CY2023Q4I.json One value per filer for that period

The CIK is ten digits with leading zeros.

The pull, in code

# sec.py
import json
import time
from pathlib import Path

import requests

HEADERS = {"User-Agent": "Your Name your.email@university.edu"}
CACHE = Path("data/sec")


def get(url: str) -> dict:
    """Fetch once, cache on disk, stay under the SEC's rate limit."""
    path = CACHE / (url.split("//")[1].replace("/", "_"))
    if path.exists():
        return json.loads(path.read_text())
    response = requests.get(url, headers=HEADERS, timeout=30)
    response.raise_for_status()
    time.sleep(0.15)
    CACHE.mkdir(parents=True, exist_ok=True)
    path.write_text(response.text)
    return response.json()


def company_facts(cik: int) -> dict:
    return get(f"https://data.sec.gov/api/xbrl/companyfacts/CIK{cik:010d}.json")


def filings(cik: int) -> list[dict]:
    recent = get(f"https://data.sec.gov/submissions/CIK{cik:010d}.json")["filings"]["recent"]
    keys = ["form", "filingDate", "accessionNumber", "primaryDocument"]
    return [dict(zip(keys, row)) for row in zip(*(recent[k] for k in keys))]


def document_url(cik: int, filing: dict) -> str:
    accession = filing["accessionNumber"].replace("-", "")
    return f"https://www.sec.gov/Archives/edgar/data/{cik}/{accession}/{filing['primaryDocument']}"

Total assets for one firm, every period it reported:

facts = company_facts(320193)  # Apple
assets = facts["facts"]["us-gaap"]["Assets"]["units"]["USD"]
print(assets[-1])  # end date, value, form, filing date, accession number

You do not have to type this. Describe it to the agent and check what it writes against the table above.

Tell the agent the rules once

Add this to the instruction file from our guide to instruction files and skills.

## SEC data
- Use sec.py for every request. Never call sec.gov directly.
- Under 10 requests a second. Cached files in data/sec/ are never fetched twice.
- The company key is the CIK. Keep it as an integer and pad to ten digits only in URLs.
- Record form type, filing date and accession number for every value you keep.
- Use the filing date, not the period end, for anything that will be matched to returns.

The last line is the one that saves a paper. A value is known to the market on the day it is filed, and an agent will not assume that unless you say so.

Three analyses worth automating

A fundamentals panelCompany facts for every firm in your sample, reshaped to firm-quarter, with the filing date attached to each number.
Text that changedRisk factors from each year's 10-K, compared with the year before, so you measure what was added and removed.
Event timestampsEvery 8-K with its filing date and item numbers, ready to be the event file in an event study.

For anything where a language model reads the text, run it on more than one model, keep a hand-labeled sample, and test for look-ahead bias.

Make it a skill

---
name: sec-panel
description: Build or refresh the firm-quarter panel from SEC XBRL data. Use when asked to update fundamentals, add firms or add concepts to the panel.
---

# SEC panel

1. Read the firm list from config/ciks.csv and the concept list from config/concepts.csv.
2. For each firm, call company_facts through sec.py.
3. Keep 10-K and 10-Q values only. Where a period was reported more than once, keep the first filing and store the later ones in restatements.parquet.
4. Write data/panel.parquet with cik, concept, period end, filing date, value, form and accession number.
5. Run pytest. Report firms with no data and concepts missing for more than 20% of firms.
6. Write a manifest with row counts and the date of the run. Do not change the firm or concept lists.

Run it unattended

Claude Code runs without a terminal session through claude -p, and its documentation notes that it “exits with code 0 on success and a non-zero code when the run fails, so your scripts can branch.”

claude -p "Run the sec-panel skill and summarize what changed since the last manifest" \
  --allowedTools "Read,Edit,Bash"

Put that line in a scheduled job and the panel is current every morning. Limit the allowed tools to what the task needs.

The hard part of SEC data is not the download. It is matching the SEC’s company key to the identifiers in your returns and accounting data. If your school subscribes to WRDS, the wrds Python library gets you there.

Start by looking at what you have. These calls are in the library’s own README.

import wrds

db = wrds.Connection()
db.list_libraries()                      # what your subscription includes
db.list_tables(library="comp")           # tables in one library
db.describe_table(library="comp", table="company")

Compustat’s company table carries the CIK next to its own key, and the CRSP link table connects that key to CRSP’s. Check both table names against your own subscription before you rely on them.

link = db.raw_sql("""
    select c.cik, c.gvkey, l.lpermno as permno, l.linkdt, l.linkenddt
    from comp.company as c
    join crsp.ccmxpf_linktable as l on l.gvkey = c.gvkey
    where c.cik is not null
      and l.linktype in ('LU', 'LC')
      and l.linkprim in ('P', 'C')
""", date_cols=["linkdt", "linkenddt"])
link["cik"] = link["cik"].astype(int)

Merge on cik, then keep rows where the filing date falls inside the link dates. Now every number and every sentence from a filing sits beside the return that followed it.

Keep the licensed side out of the model’s context. The agent writes the query and you run it, as in our guide to agents and licensed data.

References