AI lab
Three tools for the mess around the work.
What I keep meeting is never the design work itself. It’s the request that arrives as one line, the decision nobody wrote down, and the career record that only gets built the week somebody needs it.
The short version
One of these was graded under controlled conditions, by me, on invented material. One showed early usefulness inside a workplace. One never made it into delivery.
The third one is the reason the other two are written the way they are. It changed how I introduce anything I build for other people.
01
What keeps going wrong?
Three separate frustrations, all of them the same shape. Something important is known by one person at one moment, and there is no cheap way to write it down while it is still true.
-
Before the work
The one-line request
A story arrives saying the page needs a button. Design cannot tell what problem it solves, engineering cannot tell what it triggers, and testing has nothing to check it against.
-
During and after
The decision nobody kept
A design review settles something, then the reasoning evaporates. Six months later a new joiner inherits the screens without the argument that produced them.
-
Across a career
The record you build too late
Portfolios get maintained in an emergency. By then the evidence is scattered across files, tickets and memory, and hindsight quietly turns an uncertain signal into a clean outcome.
I keep trying to turn something chaotic into something calmer and more usable. These three are that habit pointed at my own practice instead of at a product.
02
How do the three fit together?
They cover one piece of work from the request to the story you tell about it years later. Each one hands off to the next.
-
01
User Story clarifies
What are we solving, for whom, and what has to be true before this is finished? It works on the request, before anyone designs.
-
02
UX Documentation preserves
What did we decide, what was only a comment, and what is still unresolved? It works on the design and delivery context, and gives the answer somewhere findable to live.
-
03
Design Evidence keeps
What did I actually contribute, how strong is the evidence, and what is still unknown? It works on a private career record, which later supports a case study, a CV line or an interview answer.
The handoff is deliberately manual. A person decides what moves from one to the next, because delivery documentation carries operational detail that should never land in a career record, and a career record carries judgements about people that have no business in a project wiki.
03
Why they are not the same size on this page
Two of them get a full page and one gets a few hundred words further down. That is not modesty about the third. It is what the evidence supports, and deciding how much room a piece of work has earned is part of the job.
-
Internally evaluated prototype
Design Evidence. Built, then graded against a rubric I wrote, using invented material. One run per evaluation.
It does not mean another designer has used it, that it has met real project evidence, or that any outcome improved.
Gets the full case study.
-
Private workplace tool
UX Documentation Assistant. It runs inside a company, I use it in my own work, and selected colleagues were given access to try it.
It does not mean a company-wide rollout, measured time saved, or documentation that improved across a team.
Gets a shorter page, because the test was smaller.
-
Private workplace prototype
User Story Agent. Built, tested on real incomplete input, and shared with stakeholders for evaluation.
It does not mean adoption. Nothing it produced entered delivery.
Gets a note, because that is the size of the evidence.
04
The two with enough behind them
-
Case study
Standardise the evidence, not the story
Read it →The screens survive a project and the reasoning does not. A privacy-first skill that captures decisions, contribution and evidence status at milestones, keeps that record private, and helps turn it into a case study, a CV line or an interview answer without making it less true.
20 of 20 against a 16 of 20 baseline One run each, never used on real evidence
-
Workplace experiment
Documentation should not need a blank-page expert
Read it →I joined a team with almost no record of its own decisions, and the reason I was given was time. I think it was the blank page. A short conversation in Teams that returns a structured draft, plus somewhere predictable to keep it, and the lesson that removing the blank page was only half of it.
One colleague, roughly five minutes, recalled No habit formed, adoption not measured
05
The one that did not progress
On a client portal project, story quality depended entirely on who wrote it. Some solution architects handed over full requirements. Other stories gave a designer almost nothing. One amounted to a request for a button: no action, no trigger, no outcome, and nothing QA could later use to tell whether what got built was what was meant.
So I built a Copilot agent for the thin ones. It establishes the story type first, then asks no more than two questions at a time, follows the branch that kind of work needs, and marks anything unresolved To be confirmed. The promise was narrow on purpose: raise the floor when a request arrives incomplete, rather than outwrite the people who were already good at this.
Fictional demonstration
Project Willow, a workshop waitlist
An invented service and an invented request, shown as it arrived and as it came back. Fixed content, nothing sent to a model.
-
What arrived
New feature: add a Join waitlist CTA when a workshop is full.
That is the whole story. It names a control and a condition. Not who can join, not what the action triggers, not how anybody finds out a place has opened.
One line -
What came back
A first draft carrying the user and the problem, expected behaviour, the joining, joined, error and already-on-the-list states, acceptance criteria, dependencies, and a definition of done that includes keyboard and screen reader review.
Five items stayed open rather than being filled in.
Reviewable draft
The part I would defend hardest sits between those two cards. Asked how somebody finds out a place has opened, the stakeholder said to tell them when a spot opens. The agent asked one follow-up, got email, and stopped. How long a place is held had not been agreed by anyone, so it stayed To be confirmed. Two more questions would have produced a number, and the number would have been invented, and invented rules in a delivery tool get treated as agreed by everybody downstream.
Project Willow is invented. The workplace project that prompted the idea does not appear here in any form.
Why it stopped there
I tested it myself on incomplete stories from real work, then shared it with the wider stakeholder group. One experienced stakeholder told me they wrote better stories than the agent did. They were probably right, and they were not who it was for, which was my error in how I framed it rather than theirs in how they read it.
Underneath that, I had shared a prototype before securing any of the things that would let it go anywhere: a committed pilot group who had agreed to try it on live stories, somebody who owned the story-writing workflow, or a route into the delivery tool. No agent-generated story entered delivery on that project. Nothing about time saved or story quality was measured, so there is no result here, only a workflow that was designed and tested.
What I would do differently is not a feature. I would position it explicitly as help for incomplete requests, then run the comparison I skipped: an original story beside the agent-assisted version, with design, engineering and QA marking both against one completeness rubric. One scenario per story type would show whether each branch surfaces the missing delivery detail without becoming heavier than writing the thing by hand.
06
What changed about how I build
Two of these worked as designed and neither changed what a team did. That is one finding rather than two, and it has changed the order I do things in.
Before I build anything for other people now, I want a name. Somebody who has agreed to try it on live work, whose work is the reason the thing exists. Without that I am building on a guess about demand and calling it a prototype.
- Find the committed pilot participant before building deeply, not after.
- Design the workflow with the people expected to use it, so it fits the work they already have.
- Put an original artefact beside an agent-assisted one and let the team judge both against the same rubric.
- Agree the adoption signal before rollout, so there is something to be wrong about.
- Introduce it in a live working session, never a link.
- Land it where the work already happens, not in a new place people have to remember.
- Treat the behaviour change as part of the product, because it is the part that fails.
None of that is really about AI. It is what shipping any internal tool asks for, and I had to build two before I took it seriously.
07
What stays private, and why
Two of these run inside a workplace. What is on these pages does not.
Every demonstration you can open here is fictional. Project Tide, Project Meadow and Project Willow are invented services with invented teams, written so the behaviour that matters can be shown without a single workplace document. The complete instructions, the real inputs and outputs, the internal tools and links, and the private repository all stay where they are.
Why rebuilt rather than redacted
Blur is not anonymisation, and removing a company name does not make a project unidentifiable. A distinctive combination of industry, feature, timing and team still points at one organisation. So the demonstrations were written from scratch in a different domain instead of being scrubbed.
There is a cost to that and it is worth naming. You cannot check my workplace evidence, so the pages carry the qualification beside every claim instead. Where something is a recollection I have said so. Where nothing was measured I have said that too.
The shipped product work is next door, and it is where the harder evidence lives.
All work →