Scope: 3 Horizon Model (#21), ISO 31000 Risk Process
(#22), Change Integration Model (#23), Understanding Learners (#25) —
the four models shipped Live · citable day-1 without prior
validation, per stakeholder decision.
Standard held to: the same bar the existing 19 frameworks and Models #20/#24 are held to on the Model card & arch page — held-out eval, external senior-credentialed coach annotation, Macro-F1 + Cohen’s κ vs. certified-coach consensus, subgroup balance within ±3% of population norms. Nothing here is a lighter bar “because they’re new” — it’s the identical bar, run on a compressed timeline.
Explicitly not in scope: re-validating the existing 19 frameworks, or Models #20 (Framework Meta-Model, schema layer, not a scorer) and #24 (I→T→E→Comb Growth Model, already flagged lowest-risk / data-adjacent to Trust trajectory).
The model card’s held-out set was built and annotated for the 19 utterance-level frameworks. Three of the four provisional models operate at a different unit of analysis, so the existing set has real gaps for this purpose:
| Model | Unit of analysis | Existing 12k set covers this? |
|---|---|---|
| 3 Horizon Model | Session or reporting period (needs multiple sessions to see energy-allocation trend) | Partially — has multi-session history per learner, but was never annotated for horizon labels |
| ISO 31000 Risk Process | Session, but requires tracking whether a raised concern was resolved within or across sessions | Partially — same gap, no risk-resolution annotation exists |
| Change Integration Model | Single session, but needs the organizational context around it (was this a crisis week or routine week?) — that context often isn’t in the transcript itself | Gap — context metadata wasn’t collected for this set |
| Understanding Learners | Per-learner, accumulated across sessions (style is a stable-ish trait, not a one-shot state) | Partially — same multi-session history exists, no style annotation exists |
Implication: we reuse the existing 12,000-session pool as the base (it already has 380 external credentialed-coach annotations and known subgroup balance), but layer a new annotation pass on top for these 4 specific label types, rather than assuming the existing labels transfer.
| Model | Primary metric | Pass threshold | Rationale for threshold |
|---|---|---|---|
| 3 Horizon Model | Macro-F1 vs. coach consensus (4-class) | ≥ 0.70 | Matches the model card’s own “slightly under” floor (e.g. Active listening 0.71) — below this is downgraded to Research preview, not removed |
| ISO 31000 Risk Process | Macro-F1 on resolved-vs-unresolved binary + Cohen’s κ | ≥ 0.70 F1, κ ≥ 0.60 | Binary classification bar set slightly below the 10-dimension multi-class average (0.81 mean) since this is a harder, cross-session judgment call |
| Change Integration Model | Macro-F1 (2-class: swift vs. deliberate) | ≥ 0.70 | Same floor; also report separately for “context metadata available” vs. “transcript-only” subsets — expect the latter to score lower, and that’s informative, not just noise |
| Understanding Learners — style mix | Mean absolute error (MAE) between model-predicted % mix and annotator-rated % mix, per style | ≤ 15 percentage points MAE | New metric type (regression, not classification) — set conservatively since this is the first cycle for this label type |
| Understanding Learners — L&D stage | Macro-F1 (5-class) | ≥ 0.65 | Slightly lower floor — 5-class problem with a real class-imbalance risk (early-stage “Identify & Align” sessions likely rarer than “Deliver & Support”) |
Subgroup balance check (all 4 models): re-run the existing ±3% population-norms balance check (language, gender across role levels, remote/hybrid/in-person mix) specifically on the new annotation subset, not just inherited from the base 12k pool — a new annotation pass can introduce new sampling skew even on an already-balanced base set.
| Week | Milestone |
|---|---|
| 0 (now) | Freeze annotation guidelines for all 4 label schemes (Section 2) — sign-off needed from whoever owns coaching-framework fidelity on your side |
| 1 | Pull/assemble organizational change-event metadata for Change Integration Model (the one real data gap flagged in Section 2c); calibrate annotator pool on all 4 schemes |
| 2–3 | Annotation pass on 1,500 session-windows × 4 models (parallel, not sequential — different annotator sub-pools can work each model concurrently) |
| 4 | Score models against fresh annotations; compute Macro-F1/κ/MAE per Section 3; run subgroup balance check |
| 4 (same week) | Go/no-go decision per model against thresholds — see Section 6 |
| 5 | Publish results: update Model card & arch page’s Evaluation
table with the 4 new rows (only for models that pass), OR downgrade
non-passing models from Live · citable back to
Research preview in Model Studio with a public note on
what’s being fixed |
This lands the checkpoint inside the “first 2–4 weeks of production data” window already promised in the rollout risk note — Week 4 is the decision point, not a soft target.
Live · citable, remove provisional tag,
publish the metrics on the Model card page like the existing 19.Live · citable but keep
provisional tag, publish metrics with the balance
caveat explicitly stated (matches how the model card already documents
known failure modes like “Silence-heavy exec styles” rather than hiding
them) — this is the honest middle state, not silently rounding to a full
pass.Live · citable to
Research preview in Model Studio, keep it running in
shadow-mode (scoring happens, but output isn’t surfaced to
coaches/managers) while the label scheme or model logic is revised, then
re-run this same eval cycle.Since all 4 are live from day 1 per the standing decision, real coaching sessions are being tagged by these models right now, before this eval exists. Recommend:
model_version + provisional=true flag (as
already recommended) so that once ground truth exists, you can
retroactively audit how the live-tagged production data
compares to what the validated model would have said — this turns the
“gap period” into free bonus eval data instead of a blind spot.Single markdown/PDF summary at Week 5, structured to match the format already on the live Model card & arch page (Dimension | Macro-F1 | Cohen’s κ | vs. certified coaches | Framework citation | Failure modes) — so it can be dropped straight into that page’s Evaluation table for models that pass, keeping the platform’s “one evaluation standard, no exceptions” credibility intact.