Practical AI, proven in use

Turn an AI experiment into evidence you can trust.

Define the baseline, compare approaches, record failures and human decisions, then return for the real outcome. Your work stays in this browser unless you choose to export it.

4
practice tracks
4
reviewed protocols
0
public outcome claims
10+
reports before a benchmark

Choose the work

Four equal paths. One evidence standard.

Compare tracks

The compounding loop

The result is not a roadmap. It is a reusable record of what worked, what failed, and why.

Better models may improve the experiment. The project history, evaluation cases, outcomes, reviewer decisions, and comparable reports remain useful.

  1. 1Frame a real job and baseline
  2. 2Compare versions on fixed evidence
  3. 3Record the outcome after use
  4. 4Sanitize and submit a proof pack
  5. 5Fork reviewed playbooks
  6. 6Publish benchmarks only at sufficient sample

Recently verified

Operational protocols, not success stories

All playbooks
ProtocolTrackVersionReview scope
Run a controlled support-triage pilotTest whether AI can reduce triage time while preserving routing accuracy and human approval.business1.0Protocol structure, privacy boundary, and measurement methodOpen →
Build portfolio proof in a weekly evidence cycleCreate verifiable proof of skill without inventing job outcomes or exposing employer information.careers1.0Evidence rubric, privacy boundary, and claim languageOpen →
Design an assessment with declared AI rolesUse AI in an assessment without hiding its role or weakening the evidence of learning.teaching1.0Protocol structure, learner privacy, accessibility, and measurement methodOpen →
Gate a RAG change with a fixed evaluation setDecide whether a RAG change is safe to ship using reproducible evidence instead of a demo impression.builders1.0Evaluation structure, version traceability, and rollback controlsOpen →

Benchmarks

Suppressed until the evidence is real.

No outcome benchmark is public today. Each cohort needs at least 10 reviewed, comparable reports and editorial approval.

  • AI-assisted support triage0/10
  • Weekly portfolio proof cycle0/10
  • Assessment with declared AI roles0/10
  • RAG change evaluation0/10
Read the method →

September field tests

Run the same protocol together.

See all challenges →