Article

The second system of record:

When AI containers outgrow your DMS

 

iStock-1390307383 inverted

This article, published originally by ILTA, explores the matter of how unstructured data can affect complex AI tools when information governance is not prioritised firm-wide.  

For two decades, the unstructured data problem took on a familiar shape. Documents piled up on network file shares, in personal drives, inside email, and across a dozen legacy systems nobody wanted to decommission. Information governance teams spent years, and in many firms, real money, bringing that sprawl under control: classifying it, applying retention, and closing it down defensibly. I have seen firms finish a cleanup like that, sign off on a project that ran two years and cost well into six figures, and feel the worst was finally behind them.

Then they adopted legal AI.

Within a year, the same sprawl had returned. Not within the old file shares this time, but inside the AI tools the firm had just bought to make its lawyers faster. Documents copied in from the document management system. Outputs parked in OneDrive for “temporary” storage that was never temporary, with volumes climbing month over month. Workspaces created and abandoned by the hundreds. One firm I have in mind generated close to 10,000 unlabeled AI containers in 12 months, with data moving freely between half a dozen tools. It had spent two years cleaning up 20 years of mess, and it rebuilt a good part of that mess in one.

This is the part that should give every firm pause. The DMS is the repository that firms have learned to govern. For most practices, it is the system of record: the place where matters live, where retention runs, and where ethical walls hold. The AI tools now spreading across firms are quietly becoming a second system of record. Almost nobody is governing them like one.

How the sprawl forms

It helps to be specific about the mechanism because the failure is not dramatic. No one sets out to build an ungoverned repository. It builds up gradually, one ordinary shortcut at a time.

Copilot-class assistants, standalone legal AI platforms, and the AI features now built into document management systems share a basic pattern. They let users upload, sync, and analyze large volumes of documents, then collaborate on the results. Every one of those actions can create a container: a workspace, a project, a vault, a chat with files attached.

Picture a routine diligence review. A lawyer pulls a few hundred documents out of the DMS and loads them into an AI tool to summarize and tag them. The tool now holds a copy of all of it. The lawyer shares the workspace with two colleagues, exports a memo, and saves it to a personal OneDrive because that is the fastest path. A revised set goes back in a week later. The matter closes. The container, the copies, and the export all remain, and none of them is the version of record, though any of them might be mistaken for it later.

Multiply that across a practice and the picture is easy to predict. Content is duplicated across several tools. Workspaces sit unowned. Copies of client material end up outside the governed estate, and the firm has no reliable way to see where any of it went.

What makes this worse than the original file-share problem is speed and invisibility. File-share sprawl took 20 years because it moved at the pace of people saving documents into folders. AI sprawl moves at the pace of automated copying and synced storage, so it compounds in months. It also hides, because IT and procurement treat these products as productivity tools rather than as repositories that hold client data. A tool does not appear on the records map. A repository does. Many firms have built repositories and filed them under “tools.”

 

These are repositories

 

That filing error is the root of the problem, so it is worth correcting.

If an environment stores client documents, retains them over time, and controls who can reach them, it is a repository. It does not matter that the vendor markets it as a workspace, that the lawyer calls it a chat, or that procurement bought it as a feature. Function decides the category, not the branding.

This is not a new idea. The records profession has long held that information has a lifecycle wherever it sits. ARMA’s recordkeeping principles and the EDRM lifecycle model both rest on the same premise: a record is governed by what it is and by the obligations that attach to it, not by the system that happens to hold it. A client document subject to a retention schedule and an ethical wall does not shed those obligations when a lawyer drops it into an AI vault. The obligations travel with the content.

The governance test for any AI environment is short. Does client information live here? If the answer is yes, the disciplines a firm applies to its system of record also apply here. There is no separate, lighter standard for content simply because it was generated by an AI tool.

iStock-2208418145

Where it breaks

 

When firms skip that test, the breakage clusters in four places. Each maps to a discipline IG teams already run on the DMS, which is exactly why its absence in the AI layer costs so much.

Retention and disposition. Content copied into an AI vault usually arrives with no retention rule attached. It is not scheduled for review or disposal, so it persists by default. The result is over-retention: copies of client material held indefinitely with no policy basis; the exact condition defensible disposition exists to prevent. The Sedona Conference made the point long ago that a firm cannot defend a deletion it cannot explain, and it cannot defensibly delete what it never tracked.

Access, ethical walls, and client restrictions. This is the failure that should concern a general counsel most. An ethical wall configured in the DMS does not follow a document into an AI workspace, nor do the outside counsel guidelines governing where a client’s data may live. A bank that bills the firm millions a year may require its work product to stay in U.S. data centers and off any tool the firm has not formally cleared. When a partner uploads a deal binder to an AI assistant her firm rolled out last quarter, the assistant can route processing through a region the client never approved, and a copy of the working memo can land in OneDrive overnight. Every one of those events is a breach the firm cannot audit, and the partner may not learn about it until the next OCG review. When a lawyer copies restricted content into an AI tool, the controls break silently, and the firm often has no record that it happened.

Ownership and auditability. Who owns an AI workspace that an associate created for a single matter? At most firms, no one does. When that associate leaves, the container stays: unlabeled, unowned, still holding client documents, with no audit trail of what went in or came out. Multiply that by the 10,000 containers from the opening example, and the size of the blind spot is clear.

Structure and metadata. AI vaults rarely carry the classification or the matter-centric structure that makes a DMS governable. Without metadata, content cannot be searched reliably, surfaced for a legal hold, tied to a client, or disposed of with confidence. It also breeds duplication, because nothing distinguishes a working copy from the authoritative one. A firm cannot govern what it cannot describe, and most of this content is undescribed.

The good news

 

If the diagnosis sounds bleak, the prognosis is not. This is a solvable problem, and in many cases, it can be solved with capabilities a firm already owns. The firms that struggle are the ones that try to clean it up after the fact. The firms that do well govern from the outset.

A few moves do most of the work. The first is to treat AI repositories as part of the governed estate rather than as a separate world. A firm already has a retention schedule, a classification scheme, and an ethical wall regime. Extend that perimeter to cover AI environments rather than inventing a parallel set of rules that no one will maintain. Governance that stops at the edge of the DMS is governance with a hole in it.

The second is to control how content enters these tools. Supported integrations that draw content from the system of record under existing controls are far safer than lawyers copying documents loose and saving outputs to personal storage. When content flows in through a governed path, its retention, its restrictions, and its provenance can travel with it. When it is copied by hand, they do not.

The third is visibility and audit. A firm should be able to see its AI containers the way it sees its matters: who created them, what they hold, whether they are still needed, and when they should close. Audit-trail monitoring and approval workflows turn an invisible sprawl into a managed inventory. It also helps to say so in policy. An AI use policy that names the firm’s repositories and states plainly that client content belongs in governed locations gives IT and IG something to enforce.

The starting point is an honest inventory. Before a firm rewrites a policy or buys a new tool, it should map where AI containers actually live across its environment, the assistants and workspaces, the DMS-integrated AI features, and the personal OneDrives that are quietly holding work product. The next pass is triage. Which of those containers hold client content? Of those that do, which have no owner and no retention rule attached? Those are the immediate priorities, because they are the ones most likely to violate something the firm has already promised a client. A firm that completes the exercise usually finds the scale uncomfortable, and the fix stops being theoretical. The conversation shifts from a debate about AI policy to a list of repositories to govern, ordered by risk.

The cost argument settles it. The firm in my opening example spent two years and more than $500,000 on its first cleanup. A second cleanup would be harder, because AI sprawl is more fragmented, more duplicative, and faster to grow than the file shares ever were. Every month a firm waits, the more the eventual bill grows. Doing it right at the outset is not the expensive option. It is the cheap one.

The choice

 

The system of record is one of a firm’s most durable assets. It is where institutional knowledge lives and where the firm’s obligations to its clients are kept. The question raised by AI is not whether to use these tools. The use cases are real: diligence, disputes analysis, drafting, research, document review, and knowledge discovery. Firms should pursue them.

The real question is whether the AI layer joins the governed estate or becomes a shadow system of record running beside it, holding the same client data under none of the same controls. At most firms, that is still a choice, because the sprawl is young. It will not stay a choice for long.

The first step is the cheapest and the most overlooked. Find out where client content is actually living across your AI tools today, before it becomes the next 20-year cleanup. A firm cannot govern, and cannot defend, what it cannot see.

Frame 482-1

About the author

 

Kandace Donovan is LegalRM's US Sales Director. Kandace leads firms in strengthening compliance, reducing risk, and driving greater control over their information assets.

To learn more, visit us on LinkedIn or view our homepage.