Imagine an employee asking the company’s AI assistant, ‘What is the notice period for supplier XYZ?’
The answer comes back immediately: a clear, precise quote from the contract. It sounds right, so nobody checks it. However, the quoted contract is an expired draft. The signed amendment that changed the notice period is stored somewhere else entirely.
This is a hypothetical scenario, but I believe similar situations are already occurring. It illustrates what I think is the most underestimated aspect of AI in ECM: a poor search result is obvious, but a poor AI response can be convincing.

AI turns content quality into a risk
Before AI, poor content quality was mostly just an annoyance. Searches took longer and users had to scroll through several versions of the same document, but they usually noticed when something looked wrong. Humans were still the filter.
With AI, however, this filter becomes less effective. The assistant reads the content, selects what it deems relevant and presents a coherent summary. Users may never see the other versions, outdated drafts or documents that contradict the answer.
In my post on information debt, I wrote about the shortcuts organizations take with their content and the price they pay later. AI is often the moment when that debt becomes payable. All those duplicates, forgotten drafts, and incomplete properties are now the raw material that your AI works with.
AI does not fix messy content. It simply reads it faster, summaries it with confidence and spreads the mess further.
What AI struggles with in typical repositories
When I look at how content is usually managed, I see five recurring problems.
Duplicates and near-duplicates. Which version is the truth? A person might know from experience. An AI assistant has no instinct for this unless the system gives it clear signals.
Outdated content. A procedure valid in 2019 sits next to the current one. Both are well written, and both look equally credible.
Missing or inconsistent metadata.. AI depends on context: the type of document, its status, its owner, the period it applies to. When the “Status” property is empty, or means different things in different departments, that context disappears. I covered part of this in The metadata trap. Metadata can fail by being too complex, and it can also fail by being unreliable.
Missing relationships. A contract, its amendments and its invoices tell one story. If they are stored as separate islands, the AI sees fragments. This is exactly the point I made in The forgotten power of relationships in ECM.
Unclear authority. Is this document a draft or approved? An internal opinion or an official policy? Without lifecycle information, an AI can’t tell the difference, and often neither can the user.
The good news: ECM discipline is AI readiness
Here’s what I find reassuring: All the features that make content manageable in an ECM also make it trustworthy for AI:
- Lifecycle and status tell the AI what is current and approved.
- Metadata tells the AI what a document is and what it applies to.
- Relationships show which documents belong together.
- Permissions determine what each user’s AI is allowed to see.
The practices we have promoted for years, such as clear metadata, defined workflows and linked objects, were not just housekeeping measures. They provide the context that transforms a collection of documents into a knowledge base that an AI can utilize responsibly.
I made a similar point in ‘ECM + AI: achieving optimal synergy’: AI does not replace a well-structured ECM; it depends on it.
The permissions blind spot
There is one more risk that deserves its own section because it takes many organizations by surprise.
For years, content that was shared too widely was protected by obscurity. A document might be accessible to more people than intended, but it would remain hidden because no one was looking for it. AI changes that. An assistant that can search and summarize everything a user has access to will surface content in seconds that nobody knew was reachable.
Therefore, before enabling AI features, check who can access your sensitive content. This includes broad groups, inherited rights, and shared links. It is much better to discover over-sharing during a controlled review than through an AI-generated response in front of the wrong person.
A practical way forward
Making all content ‘AI-ready’ can seem overwhelming, but it doesn’t have to be. Here is the approach I would recommend:
Start with one use case, rather than the entire repository
Choose an area where AI could clearly add value, such as contracts, quality procedures or customer documentation. Prepare that content first.
Define what ‘authoritative’ means
Decide which statuses, document types and validity dates identify trusted content. If your team cannot answer the question ‘What counts as the official version?’, neither can the AI.
Remove the obvious noise
Archive or delete expired, duplicate and abandoned documents. This is where the lifecycle thinking from my earlier posts comes in useful.
Fix the important metadata
Don’t try to perfect every property. Focus on the ones that users and AI rely on, such as type, status, owner, validity and relationships with related objects.
Review permissions before activation
NOT AFTER!
Keep a human in the loop at the beginning
Require AI answers to show their sources so that users can verify them. During the pilot, review the results and collect feedback on incorrect or unexpected answers.
Measure and provide feedback
When an answer is incorrect, establish the reason why. This is usually due to a content problem, such as a missing status or a duplicate. Fixing these issues improves every future answer.
AI can help with the cleanup too
This doesn’t mean that you have to have perfect content before you start. AI can help you achieve this. As I explained in M-Files AI: From Intelligent Metadata to AI Agents, it can suggest properties, classify documents, and identify similar content. This makes it a useful tool for identifying duplicates and filling in missing metadata.
A sensible approach is to work in cycles rather than gates: improve the content a little, use AI, see where it fails, then improve it again. However, it is important to have a human validate what the AI suggests, especially for properties that drive decisions.
A simple readiness check
Ask yourself the following questions about the content that AI would use:
- Can you tell whether any document is current and approved?
- Are the metadata for your key document types consistent and complete?
- Do you have an idea of how many duplicates exist of your most important content?
- Are related documents linked to each other?
- Do you know who can access your sensitive content, including via shared links?
- Can every AI answer be traced back to a source that users can verify?
- Is the quality of this content owned by someone?
If you answered “no” or “I don’t know” to several of these questions, that’s not a reason to stop. It’s a reason to start preparing!
Conclusion
AI readiness isn’t a separate project. Rather, it is what good ECM practice looks like when you finally ask your content to do more than just sit in a repository. It also provides a business reason to invest in quality work that was often postponed in the past.
The organizations that benefit most from AI won’t necessarily be those with the most advanced model. They will be the ones whose content can be trusted.
Have you tested AI features on your own content yet? I’d be interested to hear what surprised you.
