← Model Studio build brief

Fast-Follow Eval Plan

Fast-Follow Validation Plan — 4 Provisional Structural Models

Scope: 3 Horizon Model (#21), ISO 31000 Risk Process (#22), Change Integration Model (#23), Understanding Learners (#25) — the four models shipped Live · citable day-1 without prior validation, per stakeholder decision.

Standard held to: the same bar the existing 19 frameworks and Models #20/#24 are held to on the Model card & arch page — held-out eval, external senior-credentialed coach annotation, Macro-F1 + Cohen’s κ vs. certified-coach consensus, subgroup balance within ±3% of population norms. Nothing here is a lighter bar “because they’re new” — it’s the identical bar, run on a compressed timeline.

Explicitly not in scope: re-validating the existing 19 frameworks, or Models #20 (Framework Meta-Model, schema layer, not a scorer) and #24 (I→T→E→Comb Growth Model, already flagged lowest-risk / data-adjacent to Trust trajectory).


1. Why this can’t just reuse the existing 12,000-session eval set as-is

The model card’s held-out set was built and annotated for the 19 utterance-level frameworks. Three of the four provisional models operate at a different unit of analysis, so the existing set has real gaps for this purpose:

Model Unit of analysis Existing 12k set covers this?
3 Horizon Model Session or reporting period (needs multiple sessions to see energy-allocation trend) Partially — has multi-session history per learner, but was never annotated for horizon labels
ISO 31000 Risk Process Session, but requires tracking whether a raised concern was resolved within or across sessions Partially — same gap, no risk-resolution annotation exists
Change Integration Model Single session, but needs the organizational context around it (was this a crisis week or routine week?) — that context often isn’t in the transcript itself Gap — context metadata wasn’t collected for this set
Understanding Learners Per-learner, accumulated across sessions (style is a stable-ish trait, not a one-shot state) Partially — same multi-session history exists, no style annotation exists

Implication: we reuse the existing 12,000-session pool as the base (it already has 380 external credentialed-coach annotations and known subgroup balance), but layer a new annotation pass on top for these 4 specific label types, rather than assuming the existing labels transfer.

2. Ground-truth annotation design, per model

2a. 3 Horizon Model (#21)

2b. ISO 31000 Risk Process (#22)

2c. Change Integration Model (#23)

2d. Understanding Learners (#25)

3. Metrics & pass thresholds

Model Primary metric Pass threshold Rationale for threshold
3 Horizon Model Macro-F1 vs. coach consensus (4-class) ≥ 0.70 Matches the model card’s own “slightly under” floor (e.g. Active listening 0.71) — below this is downgraded to Research preview, not removed
ISO 31000 Risk Process Macro-F1 on resolved-vs-unresolved binary + Cohen’s κ ≥ 0.70 F1, κ ≥ 0.60 Binary classification bar set slightly below the 10-dimension multi-class average (0.81 mean) since this is a harder, cross-session judgment call
Change Integration Model Macro-F1 (2-class: swift vs. deliberate) ≥ 0.70 Same floor; also report separately for “context metadata available” vs. “transcript-only” subsets — expect the latter to score lower, and that’s informative, not just noise
Understanding Learners — style mix Mean absolute error (MAE) between model-predicted % mix and annotator-rated % mix, per style ≤ 15 percentage points MAE New metric type (regression, not classification) — set conservatively since this is the first cycle for this label type
Understanding Learners — L&D stage Macro-F1 (5-class) ≥ 0.65 Slightly lower floor — 5-class problem with a real class-imbalance risk (early-stage “Identify & Align” sessions likely rarer than “Deliver & Support”)

Subgroup balance check (all 4 models): re-run the existing ±3% population-norms balance check (language, gender across role levels, remote/hybrid/in-person mix) specifically on the new annotation subset, not just inherited from the base 12k pool — a new annotation pass can introduce new sampling skew even on an already-balanced base set.

4. Sample size & annotator plan

5. Timeline

Week Milestone
0 (now) Freeze annotation guidelines for all 4 label schemes (Section 2) — sign-off needed from whoever owns coaching-framework fidelity on your side
1 Pull/assemble organizational change-event metadata for Change Integration Model (the one real data gap flagged in Section 2c); calibrate annotator pool on all 4 schemes
2–3 Annotation pass on 1,500 session-windows × 4 models (parallel, not sequential — different annotator sub-pools can work each model concurrently)
4 Score models against fresh annotations; compute Macro-F1/κ/MAE per Section 3; run subgroup balance check
4 (same week) Go/no-go decision per model against thresholds — see Section 6
5 Publish results: update Model card & arch page’s Evaluation table with the 4 new rows (only for models that pass), OR downgrade non-passing models from Live · citable back to Research preview in Model Studio with a public note on what’s being fixed

This lands the checkpoint inside the “first 2–4 weeks of production data” window already promised in the rollout risk note — Week 4 is the decision point, not a soft target.

6. Go / no-go decision rule per model

7. What happens to live production data collected before this eval completes

Since all 4 are live from day 1 per the standing decision, real coaching sessions are being tagged by these models right now, before this eval exists. Recommend:

8. Reporting

Single markdown/PDF summary at Week 5, structured to match the format already on the live Model card & arch page (Dimension | Macro-F1 | Cohen’s κ | vs. certified coaches | Framework citation | Failure modes) — so it can be dropped straight into that page’s Evaluation table for models that pass, keeping the platform’s “one evaluation standard, no exceptions” credibility intact.