Start with the work, not the category
Your organisation probably has no shortage of AI ideas. Someone wants a chatbot. Somebody else has seen a demonstration that turns a meeting into a set of actions. A team is already using a public tool to draft text. A supplier has promised a surprisingly smooth future.
The difficult part is not generating ideas. It is choosing one piece of work that is useful enough to matter, ready enough to test and bounded enough to learn from safely.
"Use AI in customer service" is not a project. "Create a first draft response to routine delivery-status questions, using approved order information, for an adviser to review" is much closer.
- the trigger that starts the work;
- the people doing and receiving it;
- the information used;
- the output or decision required;
- the friction in the current approach; and
- the human judgement that must remain.
Look for repeated effort with visible inputs and outputs
Early opportunities are easier to learn from when the work happens often enough to observe. You need realistic examples, a way to compare the current approach with the test and people who can judge whether the output is useful.
Good candidates often include a defined form of finding, classifying, extracting, comparing, summarising or drafting. The task should be specific. The accountable outcome should still belong to a person.
This does not mean the most repetitive task is automatically the best one. A repeated task may depend on poor information, an unclear decision or a process that should be removed. AI can make a bad step happen faster without making the work better.
Make the consequence of error explicit
Ask what happens if the output is wrong, incomplete, biased, outdated or confidently misleading.
The answer should determine the strength of your boundaries and review. A low-consequence internal draft that an experienced person checks is different from an output that affects someone's safety, rights, employment, finances or access to a service.
If the consequence is high, the first test may need specialist legal, security, data-protection or domain advice. It may also be the wrong first project.
Check the information before blaming the model
AI work depends on the information around it. If people cannot find the current guidance, if important fields are missing or if several sources disagree, the opportunity may really be a knowledge or process problem.
Sometimes the best result of an AI opportunity review is a clearer information system. That is progress, even if no model is deployed.
- Do we have the right to use this information in this way?
- Is it current, sufficiently complete and representative of the real work?
- Does it contain personal, confidential, regulated or copyrighted material?
- Can the person reviewing the output reach the original source?
- Who owns the source and corrects it when it changes?
Name the owner and the people doing the review
A project needs an owner who cares about the work, not only the technology. That person should be able to make decisions about the workflow, involve the people affected and judge whether the result is worth maintaining.
Human review also needs a real design. "A human is in the loop" is not enough. Which human? What are they checking? Can they see the source? Do they have time and authority to challenge the output? What happens when they are uncertain?
If review requires more effort than the original task, the project may still teach you something-but it has not yet improved the work.
Choose a measure tied to the work
Avoid treating use of the tool as proof of success. Choose one or two measures that describe the job you are trying to improve.
Record the current position where you can. A test without a baseline can produce enthusiasm but very little evidence.
- elapsed time from trigger to useful finish;
- the amount of rework;
- the rate and type of errors;
- the effort required to find or check information;
- the quality and consistency judged by experienced people;
- the confidence of the person responsible for the output; or
- the experience of the person receiving it.
Make the first test deliberately small
A useful test has a clear start, owner, set of examples, review method and end date. It also has conditions for stopping or changing course.
You might test with historic or synthetic examples before any live use. You might restrict the task to one type of request. You might ask the system to draft while a qualified person remains responsible for the final output. The right shape depends on the consequence and context.
Small does not mean trivial. The work should matter enough that people can judge its value. It should simply be bounded enough that failure creates learning rather than unmanaged harm.
A simple selection check
Before choosing the project, ask whether you can say yes to these statements:
If several answers are no, do not force the idea through an AI project template. Improve the definition, information or ownership first.
- The work is specific enough to observe.
- The current friction and useful outcome are clear.
- The required information is available and appropriate to use.
- A named person owns the test and the working outcome.
- Human review matches the consequence of an error.
- Success can be judged against the work, not the novelty.
- The test is small enough to stop, change or reverse.
The first project is also a capability project
The durable value is not only the output of one test. Your people should become better at defining opportunities, judging information, designing review, spotting failure and deciding when not to use AI.
That is why the first project should leave behind a reusable canvas, examples, decisions, boundaries, measures and an owner-not just a demonstration.
A good first AI project is small enough to learn from, important enough to matter and clear enough to stop if the evidence says it should.