Analysis

AI has arrived for finance research, and data vendors cannot police it

Agents now operate a computer better than the human baseline on the standard test. Licenses forbid most AI use of data, and almost none of it can be detected.

AI Fin ResearchCovers Global

Fully automated pipelines are not a forecast. Four fields already have systems that run most of the research loop, and the last obstacle, software with no API, is gone.

12%of computer tasks the best model completed when the OSWorld benchmark launched
72%completed by humans on the same tasks
86%completed by the top agent on the OSWorld-Verified board on 9 October 2026

OSWorld tests an agent on real desktop work: browsers, spreadsheets, office files, a code editor. When it launched, its authors reported that humans completed 72.36% of tasks and the best model 12.24%. On 9 October 2026 the board had a general-purpose model at 86.0%. An agent can now use a terminal, a portal or a web query form about as well as a research assistant, and it does not get bored.

What the licenses say

Vendors and publishers saw this coming and wrote it into contracts. University libraries are passing the terms on in plain words.

Source What it says
University of Waterloo Libraries Agreements with publishers “do not allow sharing licensed materials with third parties,” including generative AI services. Breaches “can result in loss of access for individual users or the entire campus”
American University Library Contracts “explicitly forbid uploading, processing, or otherwise using the content in AI systems.” This applies to all AI tools, including the ones the university licenses

Why it cannot be policed

The contract is clear. Enforcement is another matter, for four reasons.

  1. The agent is you. An agent that works inside your own licensed session sends the requests you would send, from your machine, under your login. On the vendor’s side there is one user.
  2. Detection looks for volume. The controls vendors have are built to catch bulk downloading. Agent use does not need to look different from a heavy human user.
  3. Local models leave no trace. When the model runs on the researcher’s own machine, nothing goes to a third party and nothing is observable from outside.
  4. Only crude use leaves a mark. Early image generators gave themselves away with six-fingered hands, and sloppy AI text still does the same with stock phrases and invented citations. Competent use has no tell. A coefficient, a table or a clean paragraph does not show how it was produced.

A rule that cannot be detected gets enforced by other means.

What vendors do instead

They have already started.

Move Where it shows
Build the sanctioned door and charge for it Connectors and apps inside Claude and ChatGPT from FactSet, S&P Capital IQ, MSCI, LSEG, Moody’s and others (our map)
Tie access to identity OpenAI’s shared sign-in work with S&P Capital IQ, LSEG, MSCI, Dow Jones Factiva and Moody’s, so a provider recognizes the user’s entitlement
Price by use, not by seat The direction metered access leads. We expect it at renewal
Put the risk on the institution Campus-wide loss of access for one person’s breach, as Waterloo’s wording already says

What this means for you

Hard to detect is not the same as allowed. The license binds you whether or not anyone is watching, and the penalty lands on your whole institution. Our guide to agents and licensed data covers what you may do.

The researchers who lose here are the ones who do nothing: no agents because the rules look frightening, no conversation with the library, no say in the next contract. Their colleagues at schools that negotiated agent use will be running pipelines they are not allowed to build.

Sources

Related