Built by a tutor, for the work after every lesson

What happened today should shape what happens next.

TutorOS turns real session evidence into a focused lesson plan, a clear mastery decision, and a parent update that actually says something.

No loginNo real student dataEvidence before praise

One connected workflow

From context to an honest update.

Each output is connected to what the tutor actually observed—not an isolated AI generation.

  1. 01

    Plan the session

    Turn the last lesson and current struggle into a focused 45-minute plan.

  2. 02

    Capture evidence

    Record the moments that reveal what landed and what still needs support.

  3. 03

    Decide what is next

    Use observed performance to schedule the right topic at the right time.

  4. 04

    Shape the next lesson

    Carry the review date, support level, and source evidence into the next plan.

  5. 05

    Update the parent

    Create a warm, honest update grounded in the session—not generic praise.

  6. 06

    Tutor signs off

    Require a human review of the evidence, next step, and parent wording before use.

Fresh regression proof

12/12 evidence integrity checks passed.

12/12Measured, not claimed

Twelve synthetic adversarial fixtures run against the same mastery, Honesty Gate, and next-session functions used by the demo. No model call or API key is involved. This is a product regression benchmark—not a claim of validated learning efficacy.

Mastery scheduling4/4
Report integrity4/4
Closed-loop provenance4/4
Mastery scheduling4/4 passed
  1. Secure starts at the 80% boundary

    Expected: 80% → Secure, review in 14 days

    Observed: 80% → Secure, review in 14 days

  2. Declining attempts override the average

    Expected: Decline signal → Needs reinforcement, review in 3 days

    Observed: Decline detected → Needs reinforcement, review in 3 days

  3. A strong average cannot hide an independent miss

    Expected: 83% with independent miss → Needs reinforcement, review in 3 days

    Observed: 83% with independent miss detected → Needs reinforcement, review in 3 days

  4. Invalid session evidence is rejected

    Expected: Impossible date and empty observation → rejected before scoring

    Observed: Invalid evidence rejected by schema

Report integrity4/4 passed
  1. Invented evidence citations are blocked

    Expected: Unknown attempt ID → Honesty Gate blocked

    Observed: Blocked: The report cited evidence that is not present in this session.

  2. Generic praise cannot replace evidence

    Expected: Generic praise → Honesty Gate blocked

    Observed: Blocked: The report uses generic praise instead of a specific session detail.

  3. Non-secure mastery cannot be softened

    Expected: “On track” for non-secure evidence → Honesty Gate blocked

    Observed: Blocked: The report softens a non-secure mastery decision into misleading reassurance.

  4. A difficult attempt must be acknowledged

    Expected: Non-secure report without difficult source → Honesty Gate blocked

    Observed: Blocked: A non-secure report must cite and plainly acknowledge a difficult attempt.

Closed-loop provenance4/4 passed
  1. The latest difficulty becomes the review target

    Expected: Attempt 4 target + Attempt 3 breakthrough → carried forward verbatim

    Observed: Review target: practice-4; Breakthrough: practice-3; observations match

  2. Supported success receives a developing progression

    Expected: Modeled success → Developing; Prompt → independent → stretch

    Observed: Developing; Prompt → independent → stretch

  3. Secure evidence still keeps a traceable retrieval source

    Expected: All independent correct → Secure; latest attempt retained; 14-day review

    Observed: Secure; 14-day review; source practice-4

  4. The next lesson receives validated evidence-derived context

    Expected: Brief schema valid; API context cites Attempt 4 observation

    Observed: Schema valid; lesson context grounded in Attempt 4

Benchmark v1.0.0Fixture set 2026-07-18Executed 21 Jul 2026, 16:18 UTCpnpm benchmark
MK
Tuesday workspaceMaya · Mathematics

A tutoring session, made visible.

Synthetic student data
01

Session context

Editable input
Uses a highlighted local mock without a key · uses live GPT-5.6 when configured

A complete sample plan is ready to edit.

02

Lesson plan

Sample plan
45 minutes · 4-part planEvery section is editable
5 minReconnect with equivalent fractions

Look for: Maya explains why each pair represents the same quantity.

20 minBuild a common unit
15 minMove from models to independence
1
2
3
4
5 minExplain the method

Evidence of understanding: Maya explains that the pieces must be the same size before they can be combined.

03

Session evidence

Editable log
Attempt 1

Use fraction bars to solve 1/3 + 1/6.

Attempt 2

Solve 1/4 + 1/6 using a shared multiples list.

Attempt 3

Solve 2/3 + 1/4 without a visual model.

Attempt 4

A recipe uses 3/4 cup oats and 2/3 cup nuts. How much altogether?

Synthetic evidence is preloaded. Edit it, then update the decision.

04

Mastery decision

Review
60%evidence
Unlike denominatorsNeeds reinforcement

An independent attempt was incorrect, so this topic returns in 3 days instead of relying on the 60% average alone.

Review on 2026-07-17

3-day interval · transparent TutorOS heuristic

05

Three-session learner trajectory

Current
DirectionRecent transfer gap

Independent success appeared across the three sessions, but Tuesday · Transfer check ended with an independent miss. The latest evidence keeps the topic in a 3-day review cycle.

  1. 2026-06-30Session 1 · Build the model
    45%
    Needs reinforcement3-day review
    Independent success
    0/2
    Support used
    2/2
  2. 2026-07-07Session 2 · Fade the prompt
    63%
    Developing7-day review
    Independent success
    0/3
    Support used
    2/3
  3. 2026-07-14Tuesday · Transfer check
    60%
    Needs reinforcement3-day review
    Independent success
    1/4
    Support used
    2/4
06

Next session brief

Closed loop
Scheduled review2026-07-17
Support progressionModel → prompt → independent
TargetUnlike denominators
5 min
Retrieve Unlike denominators

Rework “A recipe uses 3/4 cup oats and 2/3 cup nuts. How much altogether?” with a visual model, then remove the model for one fresh example.

Look for: Maya completes the fresh example without repeating the recorded difficulty.
Move
Build independent accuracy in Adding fractions with unlike denominators

Restore the visual model, fade to one prompt, then check the same idea independently.

Why now: An independent attempt was incorrect, so this topic returns in 3 days instead of relying on the 60% average alone.
Check
Independent transfer

Solve one new unlike denominators problem independently and explain each step.

Look for: Independent completion without the support used at the start of the session.
Evidence carried forward
  • Review target · Attempt 4Added denominators directly after the visual model was removed.
  • Breakthrough · Attempt 3Found 12 as a common denominator independently.
Local mock without a key · live GPT-5.6 when configured · deterministic decisions stay unchanged

The next teaching move is ready without a model call.

07

Parent update

Honesty Gate passed
Safe synthetic sample
Uses a highlighted local mock without a key · uses live GPT-5.6 when configured

A safe sample is ready. Generate again after changing the session evidence.

Honesty Gate: passed

Passed — this report cites recorded session evidence and preserves the mastery decision.

  • Attempt 3Found 12 as a common denominator independently.
  • Attempt 4Added denominators directly after the visual model was removed.
08

Tutor sign-off

Review required
Human review remains the final boundary.

Sign-off revalidates the current evidence, mastery decision, next-session sources, and parent wording. Any later edit revokes it.

Review the current evidence, next-session brief, and parent wording before sign-off.