Approved sources for AI: how to choose them
7 minutes read

The open web is not a source, it is everything at once. Approved sources are the specific material a team has chosen to trust, and choosing them well is most of the work.
An approved source is external material a team has deliberately chosen to trust for AI-assisted work: a specific report, standard, or reference site someone reviewed and added on purpose. The open web is the opposite of that. It is not a source at all, it is every source at once, good and bad, current and stale, with no one having decided any of it belongs in this firm's work.
That distinction sounds small until a draft ships with a claim traced back to a content farm article that happened to rank for the right search term. Approved sources exist to keep that from happening, and building a good set is less about finding more material and more about deciding, in advance, what counts.
Why an AI default to the open web causes problems
The open web fails as a default source because nothing about it was chosen for this firm's work. A general web search returns whatever ranks, and ranking rewards volume and recency signals, not accuracy or relevance to a specific engagement. A model drafting from that pool has no way to tell a regulator's own guidance page from a blog post summarizing it three steps removed, with details drifting further from the original at each step.
The failure is quiet at first. A first draft built this way still reads fluently, because fluency was never the problem. What breaks is traceability: when a reviewer asks where a claim came from, "the model searched the web" is not an answer anyone can check in the time review is supposed to take. The reviewer either trusts it or starts the research over, and starting over is slower than if the draft had never claimed a source at all.
What makes a source approved, not just available
A source earns a place in an approved set on a few specific grounds, not because it turned up first.
| Criterion | What it means | Example |
|---|---|---|
| Traceable | Someone can open the exact document and check it | A named report or standard, not a search result |
| Chosen for this work | Reviewed and added for a specific kind of task | The regulator's own guidance, not a summary of it |
| Access-controlled | Only what the firm added is in scope by default | A maintained list, not an open web search |
| Kept current | Someone owns refreshing it before it goes stale | A quarterly check on anything tied to a live standard |
The second row is the one teams skip most often. It is tempting to approve a source because it looks credible in isolation, without asking whether it fits the kind of work it will get used for. A well-regarded industry outlet is a fine source for a market trend and a poor one for a legal obligation. The list should be built task by task, not as one generic "trusted sites" folder.
A worked example
Say a firm is drafting a compliance summary for a client, and the underlying question is what a specific regulation requires. An approved source for this is the regulator's own published guidance: a document with a version number and a date, something the reviewer can open directly and match a sentence in the draft against.
An unapproved default would be a general web search for the regulation's name, which returns a mix of the regulator's page, three law firms summarizing it with their own framing, and a forum post from two years ago that has not been updated since an amendment. A model pulling from that mix has no way to weight the regulator's page above the outdated forum post. Both look equally citable in the search results, and the draft inherits whichever one shaped the sentence first.
The fix is not smarter search. It is deciding, before the draft starts, that this kind of work draws only from the regulator's own material, and treating anything else as something the reviewer has to bring in deliberately, not something the model finds on its own.
How to build and maintain an approved source set
Start with what the team already checks by hand today, the reports, standards, and reference sites an experienced person on the team would open without thinking. That list is usually shorter than it feels, and it is the right starting point because it reflects judgement the firm already trusts.
- Write down the sources a senior person on the team relies on for this kind of work, not a generic "good sites" list.
- Group them by the kind of work they support. A source that is right for market context is often wrong for a compliance claim.
- Assign an owner for each group, someone whose job includes noticing when a source goes stale or a newer edition replaces it.
- Set the default for that kind of work to draft only from the approved set, and make pulling in anything else a visible, deliberate step rather than something the model does quietly.
- Review the set on a fixed cadence. A source that was current a year ago may not be now, especially for anything tied to a regulation or a fast-moving market.
None of this requires an engineer. Deciding which sources the firm trusts is judgement work the person who owns delivery already does informally. Writing it down and keeping the AI inside it is what turns that judgement into something repeatable.
Where approved sources fit in the rest of the work
Approved sources are the input side of a loop that continues into review. Choosing the right sources reduces how often a claim needs to be questioned later, but it does not remove the need to check: even a claim from an approved source can be stretched beyond what the source says. The guide to telling a sourced claim from an assumption covers that check in detail, and it assumes the source itself was already a reasonable one to draw from, which is exactly what this list decides in advance.
Approved sources are also one piece of a firm's connected knowledge, alongside its own prior work, its methods, and permitted client material. The connected knowledge guide covers how all four fit together, and how they connect to the tools and data a firm already uses under the access rules those tools already have.
The short version
An approved source is something a team chose on purpose, for a specific kind of work, and can point to directly. The open web is not a substitute for that list, no matter how good any single result looks in isolation. Building the set is front-loaded work: a firm that spends an hour deciding what counts spends far less time later tracing a claim back to something nobody would have chosen if they had been asked first.
