Compare the boundary, not the brand.
Each lab starts with one architectural question, shows the project evidence state, and keeps uninspected separate from absent.
Tool Contracts
When does an API, MCP resource, shell command, or UI action become a model-usable tool?
Governed Memory
Who owns saved agent context, and how can a reader tell whether it is current, scoped, and safe?
Delegation Receipts
How does a parent agent prove what a child did, received, and returned?
Approval Boundaries
Where does consent live when an agent uses tools, wakes later, or delegates work?
Human Workbench
Can the operator see enough of the agent system to steer, review, stop, and resume work?
Sessions And Recovery
What survives interruption, restart, protocol change, or a client disappearing?
Evidence And Evaluation
What would prove that the agent did the intended work rather than merely reporting success?
Maturity Signals
Which changes show agent projects leaving prototype mode and becoming maintainable infrastructure?