Guide

Agent-first coding: use the frontier tools and pay for them

Claude Code or OpenAI Codex, desktop app or command line. Generic and budget setups hide what the new tools can do.

AI Fin ResearchStarter4 min read
Use the agent built by the lab that trains the model, on its best model, on a plan that does not run out.

What agent-first means

Editor-first Agent-first
You Type the code Describe the outcome
The tool Completes your line Reads the repository, edits files, runs commands
You check Each line as you write it The diff and the output
Good for Small edits Pipelines whose result you can verify

One session looks like this.

1. You ask"Pull CRSP monthly returns for 2000 to 2023 and check for duplicate permno-date rows."
2. The agent worksIt writes the query and the script, runs them, reads the error, fixes it and prints the row counts.
3. You reviewYou read the diff and the row counts, then ask for the next step.

Research code suits this. It is scripts and pipelines run from a command line, and the questions have checkable answers: did the pull return the right rows, does the regression reproduce the table.

“Claude Code or VS Code” is the wrong comparison. VS Code’s own documentation now describes agents that “find relevant code, make changes, and run checks without you directing each search, edit, and test run.” The editor is one of several places an agent runs. The question is which agent.

Stick to the frontier

Today that means Claude Code or OpenAI Codex, as a desktop app or on the command line.

Claude Code OpenAI Codex
Made by Anthropic OpenAI
Runs in Terminal, IDE, desktop app, browser ChatGPT desktop app, command line, IDE extension, cloud
Gets first Anthropic’s finance agent templates, as plugins OpenAI’s research tooling, bundled in its academic program
New ability lands there firstAnthropic's ten finance agent templates shipped in May 2026 as plugins for its own products. Third-party tools get features later or never.
Generic means lowest commonA tool that must work with every model cannot lean on what the best one does well. Neutral feels prudent and costs you the hard tasks.
Cheap is a different resultOpenAI's October 2026 math results took roughly three hours of top-tier thinking each. A capped plan on a small model never gets there.
A car is not a faster horse. A researcher who tests AI on a free tier is timing the horse.

The cost of economizing is invisible. You never see the analysis you did not attempt, so you conclude the technology is modest.

When you cannot use them

Some schools restrict which vendors may receive code or data, and some data cannot leave your machine at all. Then use OpenCode, an open source agent that lets you choose the model provider, or a local model. Our OpenRouter and Bedrock comparison covers where that access comes from.

Treat this as the exception. Expect less from it, and do not judge the technology by it.

The layers you can delete

Two years of tooling grew up around weaker models: orchestration frameworks, output parsers, retry wrappers, routers. Much of it is now optional.

Anthropic’s engineering team wrote in December 2024 that the most successful teams they worked with “weren’t using complex frameworks or specialized libraries.” Their advice: find “the simplest solution possible,” because frameworks “often create extra layers of abstraction that can obscure the underlying prompts and responses, making them harder to debug.” They suggest calling the model API directly.

The author of 12-Factor Agents, a guide from HumanLayer, reports the same thing from the field: “I don’t see a lot of frameworks in production customer-facing agents.” Most of what the author sees is “mostly deterministic code, with LLM steps sprinkled in at just the right points.”

For a research pipeline:

Layer Verdict Why
Orchestration framework Delete A loop in a script is easier to read, debug and put in a replication package
Parse-and-retry wrapper around JSON Delete Model APIs now offer structured outputs that, in Anthropic’s words, “guarantee schema-compliant responses through constrained decoding”
The output schema Keep You still have to say which fields you want. Anthropic’s own Python example writes it with Pydantic
Checks where outside data enters Keep Row counts, key uniqueness, date bounds. An agent removing layers should never remove these

So “you don’t need Pydantic” is half right. The retry machinery built around it can go. A typed schema is still the clearest way to state the output.

The test for any layer: if you cannot say what it does that twenty lines of your own code would not, delete it and see what breaks.

References