The work started running itself
The last update ended with a request: keep the reports coming, including the vague ones.
You did.
When we started planning v1.0.21, an ADO sweep found 511 open items. That number was a problem, and not because the product was getting worse. It was because the way we built software could no longer keep up with the way the product was being used.
So the biggest change since August isn’t in ProjectXL.
It’s in how ProjectXL gets made.
Two releases have shipped since then: 1.0.20 on September 18 and 1.0.21 today. Seventy-nine more governed build packets have closed, and the history has grown by more than 2,800 commits.
1.0.21 is the most robust release we have shipped.
This entry explains why.
The bottleneck was me
Through August, every build followed the same governed loop. A packet scoped the change, an independent review checked it against its design record, the code was implemented, and then it was walked in Excel before it could close.
Every step except the last one could run in parallel.
The walk could not.
I had to sit at a real Excel window, run a real workbook through the change, and decide whether it did what the design said. Each walk took a block of my time, and there was always a queue of builds ready for one.
The process was sound.
It had a single-threaded human in the middle of it.
Programmes, not packets
The first change was to stop treating builds as individual tickets.
Related work is now organised into programmes. A programme is a named body of work with its own charter, an ordered ladder of builds, and a running journal of every decision made along the way.
Each programme is driven by a conductor: a long-running agent session that dispatches builds, sequences reviews, maintains the journal, and brings decisions to me rather than guessing at them.
Several programmes can run at once. In late September there were usually three or four live, working different fronts: change admission, the alpha backlog, walk-backlog clearance, and structural decomposition of the shell layer.
They share one codebase and, in practice, more than one machine. They coordinate much like a development team would, with explicit ownership of files and fences around each other’s work. A single build lock decides who may compile and test at any given moment.
My role changed with it.
I stopped being the person who does each step and became the person who decides.
The v1.0.21 programme alone recorded 52 numbered operator rulings. Each was a short, specific question brought to me with the evidence already gathered and a recommendation attached:
-
Should a negative non-labor amount be legal? Yes. Refunds exist.
-
When you switch lenses, should the selected activity come with you? Yes, wherever the target lens shows it.
-
Should Delete always confirm? Yes.
Today, 18 completed programmes were formally retired into the record.
That was a short list in August.
Machines walk first
The second change is the one that broke the bottleneck.
The development environment now drives Excel itself.
A walk begins as a plan of cases with expected outcomes taken from the design note, not from someone’s memory of it. An agent launches Excel against a copy of a test workbook and drives the real add-in: real clicks into the real WebView2 surfaces and real cell edits through COM.
It then reads the result three ways:
-
what appeared on screen,
-
what ended up in the workbook, and
-
what the host logged.
Every case is scored PASS, FAIL, or INCONCLUSIVE, with evidence attached.
The rule is simple:
Drive first. Offer the human check only after the driven run looks right.
The first machine-driven walk ran on September 27. Fifteen cases took 6.5 minutes, against an estimated 25 minutes for a hand walk. Fourteen passed, none failed, and one came back inconclusive — honestly flagged rather than rounded up.
A redrive of three fixes during the final days of 1.0.21 took four minutes of wall time, including evidence capture.
But speed isn’t the main gain.
What changed is what reaches me.
By the time I sit down at Excel, the mechanical questions are already answered: Did the row count drop? Did the log line appear? Did the value persist?
What’s left is the part that needs a person:
Does this feel right? Is this what a planner expects?
On a single day this week, 50 alpha-reported items were closed in ADO, each with walk evidence attached.
Under the old process, that would have been weeks of my evenings.
What that bought: v1.0.21
v1.0.21 was planned alpha-first.
Items reported by testers went to the top of the list, ahead of internal work, and the release closed when its walks did, not when a date arrived.
Some of what changed:
-
Nothing moves under you. Feedback banners in the detail panel no longer push the controls you were about to click.
-
Selection follows you. Switch lenses with an activity selected, and it stays selected wherever the new lens shows it. In WP Timing, its work package expands and scrolls into view.
-
Delete always confirms. A successful delete also tells you it succeeded in the footer.
-
Refusals clear themselves. If an action is refused, the warning clears the next time that same action succeeds. You no longer have to wonder whether a stale warning still applies.
-
Refunds are real. Non-labor plan amounts and quantities can now be negative.
-
No identifiers in messages. Errors name the rate stack or sheet involved, not an internal ID.
-
Licensing tells the truth. An expired licence now says it has ended and explains how to restore access instead of reporting a generic verification failure.
v1.0.20 had already cleaned up the schedule network. Links no longer route through unrelated nodes. A one-day duration edit no longer reshuffles the entire graph. Link mode now highlights only the nodes you can actually connect to — in a representative plan, the old behavior highlighted 153 of 158 nodes, which told you almost nothing.
It also removed some quieter costs: 3 to 7 seconds of ribbon refresh on every workspace refresh, and a full labor recompute every time you switched lenses.
Recompute only what changed
Last time, “finish the change-admission work” was on the what’s-next list.
Most of it has landed.
The principle is easy to say and difficult to enforce:
Work runs only when something has declared that it needs that work.
An edit to one assignment should not mark every assignment in the plan as stale. A lens switch should not rebuild data nobody is looking at. A write the add-in makes to itself should not trigger another round of invalidation.
Each of those had been happening.
And each is exactly the kind of cost that gets worse as projects get larger.
On a 20,000-row test case, the whole-population freshness sweep now costs about 0.3 seconds, and individual interactions are held to a 150 ms target.
Change admission remains a live programme after this release. There is still more to narrow.
The part nobody sees, revisited
In August, the automated suite ran about 7,400 tests.
It now runs more than 12,000, and the final v1.0.21 builds closed green against it.
The machine walks didn’t replace that suite. They sit on top of it.
Unit tests prove the code does what the code says.
Walks prove the product does what the design says — inside the real Excel, with the real add-in loaded.
Before September, only one person could run that second kind of check.
Now the environment runs it, and I spend my time on the judgement calls.
Where this leaves us
The July update said the architecture had started enabling the product.
August showed that architecture could survive contact with real users.
September showed something else:
the process could scale too.
Once orchestration took the coordination off my desk and machine-driven walks took care of the mechanical verification, the limit on how fast ProjectXL could improve stopped being my calendar.
You can see the result in the release.
v1.0.21 is robust in a way the earlier alpha builds were not. The kinds of issues that made testers hesitate — unclear saves, stale warnings, controls that moved, surfaces that slowed down as projects grew — have been worked through systematically and walked against their designs.
It is the first build I’d hand to someone new without a list of caveats.
That changes what comes next.
What’s next
-
Expand the alpha. v1.0.21 is ready for more testers, and we’re getting ready to bring them on. If you know a PMI-trained planner who would put a real project through it, introduce us.
-
Start Release 22, alpha-first again. New testers bring new projects, and new projects find new edges. The machinery built this month exists to absorb exactly that.
-
Finish change admission, so the cost of an interaction stays bounded as projects grow.
-
Point the drive-and-capture harness at in-product learning. The machinery that walks a build can also record a live surface. That’s the raw material for lessons that teach against the real product.
-
Then Beta.
If you’re running the alpha, keep the reports coming.
The vague ones are still welcome.
They now reach a machine-driven walk faster than they used to reach me.