The most popular first AI project in this region is a system that searches your own documents and answers questions from them. It is popular for good reasons. It is also the project where I most often find that nobody has thought about who is doing the asking.
The failure is never dramatic. Somebody in one department asks a reasonable question, gets a reasonable answer, and the answer contains a figure from a document they would have been refused if they had gone looking for it themselves.
The system is not being disloyal, it is being obedient
Whoever set it up decided what it was allowed to read. If the instruction was the whole shared drive, then the whole shared drive is what it answers from, including the folder marked private, the board papers, the letter about somebody's conduct, and the file with everybody's pay in it.
When a person clicks into a folder they are not allowed into, the computer refuses them. The system that answers questions has already read the folder. There is nothing left to refuse.
Permission has to be checked at the moment of answering
The common shortcut is to build it on everything, then add a rule telling it not to discuss certain subjects. That is a request rather than a lock. People find their way around it within a week, usually without meaning to, by approaching the same information from the side.
What you want instead is easy to describe and does take work to build. Before the system looks for anything, it establishes who is asking, and then it searches only what that person could open by hand today. Same person, same documents, faster. Anything they would be refused in the ordinary way never enters the answer, because the system never reads it for that question.
Ask your supplier to explain in one sentence how the system decides what it may read for a given person. If the answer is about instructions given to the model, it is the wrong answer.
Your permissions are not what you think they are
This is where the project stops being about AI. Most organisations have permissions that grew rather than permissions that were designed. A folder shared with somebody in 2019 who has since changed department. A mailbox three people can open. A finance drive left open to everyone in the office because that was easier the year it was set up.
Building this kind of system exposes all of it at once. That is not a fault of the system, and in my experience it is the most useful thing the project produces, whatever else it goes on to achieve.
Do the work before you build. Take six or eight real people from different levels and different departments and write down, for each of them, what they should be able to see and what they should not. An afternoon with the right people in the room, and you have the specification for the whole permissions question.
Test it with the people who should be refused
Every test plan I am shown checks that the right person gets the right answer. Almost none check that the wrong person gets nothing. They are separate tests and only one of them ever gets run.
Sit down with somebody junior, from a department with no business seeing any of this, and ask them to get it out of the system. Not politely. The questions that break these systems are the indirect ones.
- Who is the highest paid person in this department
- Summarise what has been discussed about the restructure
- What are the payment terms in the contract with our largest customer
- Has anybody here been given a written warning this year
- What did the board decide about the new office
- Show me everything that mentions my own name
A system that refuses the direct question and answers the sideways one has failed, and on a first attempt that is the ordinary result rather than the surprising one.
Keep some things out of it altogether
Permissions are one control. Deciding what goes into the pile in the first place is another, and it is cheaper. HR case files, medical information, legal advice, anything written about a named individual: leave them out entirely unless there is a specific reason to include them. A document that was never read into the system cannot be exposed by a mistake in a rule.
Alongside that, keep a record of every question asked and every answer given, with the name of whoever asked. You will want it on the day somebody tells you the system said something, and you will want it again when you need to work out how a piece of information actually travelled.
None of this is difficult and all of it is much harder to add later, because by then the system has been built on the assumption that it may read everything. Work out who should see what before anybody points software at your documents. If you already have one running and nobody has tried it from the wrong side, that test is worth an afternoon of somebody's time this month, and if you would rather it was done by somebody from outside the organisation, ask.
Working through this on a live programme?
A 45-minute call with the engineer who would run the work. We will tell you whether AI is the answer, including when it is not.