Guide

What you may do with an AI agent on licensed data

Most data licenses forbid putting the content into AI systems. Here is what that covers, what it leaves open, and how to keep working without risking your campus's access.

AI Fin ResearchStarter3 min read
We know what the licenses say. We also know what is happening: researchers are using these tools anyway, because the fear of being left behind is real.

At a major conference in 2026, 22.5% of reviewers who were told not to use an LLM said they used one anyway. The rules are already losing. This guide is how to get the speed without betting your campus’s access on it.

What is usually allowed

This table is built from the published guidance of two university libraries, one in Canada and one in the United States. Use it to know what to ask, then ask your own library.

What you want to do Usually allowed? Why
Paste licensed data or a licensed article into a chatbot No Both libraries say licenses forbid it
Do the same in an AI tool your university pays for No American University says the restriction covers “all AI tools,” including university-licensed ones
Use the AI search or assistant built into a library database Yes American University says embedded tools “are designed to work within the licensing framework”
Use open access or public domain material with any AI tool Yes Not governed by the license
Run your own code over licensed data on your own machine Depends Waterloo treats text and data mining as a separate question from generative AI. Check the license
Run a local model over licensed data on your own machine Unsettled Nothing leaves your machine, but “using the content in AI systems” may still cover it. Get an answer in writing

Waterloo also notes where the answer is recorded: when a publisher permits generative AI use, the library says it will be noted in its catalogue. Look for the equivalent at your school before you email anyone.

How to keep working

The restriction is on the content reaching the model. Most of what an agent does for you does not need the content.

  1. Give the agent the shape of the data, not the data. Table names, column names, types, and a few rows you made up. That is enough for it to write the query and the cleaning code.
  2. Let it write code that you run. The agent writes pull.py. You run it. The licensed rows go from the vendor to your disk and never enter the model’s context. Our data pull guide is built this way.
  3. Return summaries, not rows. Have scripts print row counts, date ranges and test results. The agent can debug from a manifest.
  4. Keep outputs that contain licensed values out of the chat. A regression table of coefficients is yours. A printout of raw vendor fields is not.
  5. Write the boundary into the instruction file. One line in the repository that says which folders the agent must never read or print. It reads that file on every run.

Questions to put to your library

  • Does our license for this database permit use with generative AI? Where is that recorded?
  • Does the answer change for a model that runs on my own machine?
  • Does it change for a cloud service where the model provider cannot see prompts?
  • Is text and data mining covered separately?
  • Who signs off if I need an exception for a project?

Ask in writing and keep the reply. If the answer is no, ask what it would take at the next renewal, and tell your department head you asked. Licenses are renegotiated. Schools that ask for agent use get it sooner than schools that stay quiet.

What to record

For each project, keep a short note: which datasets, under which license terms, which tools touched what. A referee or a vendor may ask one day. The researchers who can answer in a paragraph are the ones who will keep their access.

References