Labs
The canonical software-company template uses a lab package to organize evaluation and promotion. Other companies may select a different evidence discipline; labs are not common runtime anatomy.
What a lab is
In this package, a lab is a durable environment for controlled learning — the place where the company organizes questions, experiments, specimens, instruments, evidence, results, and warrants around a target of improvement. It’s not a single experiment, eval, or test; it’s the place where proposed changes are handled as claims that must earn their evidence before production adoption.
A company using this template treats labs as a required discipline inside the selected package, not as optional ceremony. The package itself remains optional. Its purpose is to keep self-modification evidence-based instead of promoting prompts, skills, workflows, or other changes because they merely sound useful.
The pieces of a lab
- Specimens — Test cases. Real transcripts, customer conversations, synthetic cases, historical failures — the inputs you’re testing against.
- Levers — The things the company might change. Prompts, skills, models, workflows, triggers, session policies, harness assignments. Each lever can be independently varied.
- Evals — Instruments. Each eval has a range, a calibration, known noise, and known blind spots. Evals measure specific behavior — they don’t measure truth.
- Baselines — The comparison standard. Current production, a simple alternative, the previous model, doing nothing. Without a baseline, you mistake attractiveness for improvement.
- Warrants — Narrow conclusions, scoped and time-bounded. A warrant says: this specific conclusion is supported by this evidence, under these limits, for this scope, until this expiry or review condition. Warrants are how evidence becomes adoption.
Common lab shapes
- Organization-dynamics labs — Test topology, attention patterns, workflows, cognition, learning speed. Should this team be split? Does this workflow actually improve decision quality?
- Competitor-benchmark labs — External market and capability tests. How does the company’s output compare against alternatives a customer might choose?
- Customer-engagement labs — Retention, trust, channel quality. Do changes to outbound communication actually move what we care about?
- Skill-adoption labs — Test community capabilities against the company’s own use cases. Will this new skill add real capability, or just complexity?
How evidence flows
In this package a lab produces results. Results that pass review become warrants. Warrants that pass approval become promotions — into a team’s skill set, into company doctrine, into a workflow version, into a harness assignment, or into a product change. Results that fail become archived negative evidence (with date, model, harness, version, cases, assumptions, failure mode) so future agents don’t repeat the same attractive mistake.
What you operate on as the human
You propose which questions become experiments. You understand eval scope and blind spots before trusting results. You review promotion paths — where does a result go? When risk is legal, ethical, or brand-related, the company surfaces results to you for human judgment before production adoption.