Skip to content

docs(examples): example 40 must not call answer_confidence calibrated - #972

Merged
NandhaKishorM merged 2 commits into
NandhaKishorM:mainfrom
aashish254:fix-example40-calibrated-claim
Oct 7, 2026
Merged

NandhaKishorM merged 2 commits into
NandhaKishorM:mainfrom
aashish254:fix-example40-calibrated-claim

Conversation

@aashish254

Copy link
Copy Markdown
Contributor

What

examples/40_caching_and_monitoring.py calls answer_confidence "the calibrated probability of the answer Laya reports" — in its banner, in bucket()'s docstring, and in the closing paragraph. This PR makes that claim conditional and adds a Part C that prints the temperature this checkpoint actually applies to each question shape the page asks about.

The library now reads, per shape:

   question            type    k    bucket        applied   source
   intent              choice  6    choice:6-10   1.0000    bucket map
   is_urgent           noul    2    noul:2        1.9834    bucket map
   frustration         score   4    score:3-5     1.2514    bucket map
   refund_requested    noul    2    noul:2        1.9834    bucket map
   churn_risk          noul    2    noul:2        1.9834    bucket map
   shipped but refused by the runtime: choice:11+=0.1006 -> 0.5000
   at a temperature indistinguishable from 1.0: `intent` (1 of 5 question shapes)

Four pure helpers back the table: option_count, scale_for (a replay of Agent.predict_batch's own temperature_by_options.get(temp_bucket(...), temperature[QTYPES[...]]) lookup, read off the loaded agent), clamped_buckets (the condition behind the loader's RuntimeWarning), and entropy_confidence (the page's hand recomputation of confidence, printed beside the value the agent reports rather than instead of it).

Why

The word "calibrated" here is a claim about the temperatures, not about the field. laya/common.py:answer_confidence reports max(p) and its own docstring plus laya/confidence.py:17-22 say the calibration reading holds only after temperatures have been fitted and validated on held-out data. The README's Calibration section says the shipped checkpoints are over-confident and that laya-multilingual has no fitted temperatures at all. This example loads the English checkpoint and the loader emits at load time:

RuntimeWarning: laya: this checkpoint ships invalid temperatures or values outside [0.5, 5.0];
using choice:11+=0.10058280825614929 -> 0.5. Treat confidence from the affected entries as uncalibrated.

Two facts about the shipped English checkpoint drive the fix. choice:6-10 — which the page's own 6-option intent question hits — ships 1.0000158548355103, i.e. effectively unscaled: max(p) on those answers is the raw softmax. And choice:11+ ships outside [TEMP_MIN, TEMP_MAX] and is clamped. Neither is a problem the page can see without reading them back off the agent, so it printed "calibrated" without proving it.

Example 41 builds directly on this monitor and turns it into a service policy. Teaching the gate on an unqualified "calibrated" is what the follow-on deployment inherits.

The old closing paragraph computed entropy = -sum(...) and printed 1 - entropy under the label confidence, substituting its own arithmetic for the field the agent returned. The identity holds here — the vague ticket's intent answer gives 0.1404 recomputed against 0.1404 reported, 0.000014 apart — but the page never proved it. The new version prints the reported field beside the arithmetic so the reader sees they agree instead of trusting they will.

How it was verified

# CI path (script mode — this file has a module-level sys.exit like every other example gate)
python tests/test_confidence.py
# → 141 passed, 0 failed   (page-28 section unchanged: 8 helpers + 8 tests)
# → page-40 section: 7 tests

# The example itself runs end-to-end against the real English checkpoint offline:
python examples/40_caching_and_monitoring.py
# → exit=0. Prints the Part C table above and:
#   vague ticket's intent answer:  recomputed 0.1404 vs reported 0.1404, 0.000014 apart;
#   answer_confidence = 0.36

# Lint + compile gates from AGENTS.md
ruff check laya/ examples/ tests/ --select=E9,F63,F7,F82,F401,F811 --line-length=120
# → 16 F401/F811 findings across other files, none in `examples/40_caching_and_monitoring.py`
#   or my new section in `tests/test_confidence.py`. Drive-by fixes deliberately not in this PR.
python -m compileall -q laya/ tests/ examples/
# → exit=0

# Mutation harness at laya-bench/mutate_example40_gate.py, seven mutants on the example file,
# each byte-restored, run with PYTHONDONTWRITEBYTECODE=1 and PYTHONPATH pinned to the worktree:
#   M1 scale_for swaps in_map on the mapped arm      → test_scale_for_replays_core_lookup
#   M2 clamped_buckets inverts its predicate         → test_clamped_buckets_names_only_...
#   M3 entropy_confidence drops the `1 -`           → test_entropy_confidence_recomputes_core_exactly
#   M4 option_count gives noul 3 options            → test_option_count_matches_the_shipped_preset
#   M5 banner calls the probability calibrated      → test_page40_drops_the_unconditional_calibration_claim
#   M6 scale_for hardcodes "choice:6-10"            → test_page40_helpers_are_live_defer_to_core
#   M7 page stops reading the reported `confidence` → test_page40_reads_the_scaling_and_the_reported_field
# → 7/7 caught, tree restored: True

The gate execs only the example's four new top-level helper functions from its AST and injects temp_bucket, QTYPES and math from laya.common, so no weights load and the gate proves the example defers to core's real temp_bucket rather than restating a table. Each ban is witnessed against an OLD_PAGE_40 literal of main's exact wording — no ban is a guess.

Checklist

  • Focused on one change (split unrelated work into another PR)
  • Rebased on the latest main (a4a8921 release: 0.3.28)
  • Tests pass locally
  • Docs or examples updated when the public API changed

The banner, `bucket()`'s docstring and the closing paragraph all called
`answer_confidence` "the calibrated probability". `answer_confidence` is
`max(p)`, the quantity temperature scaling fits and ECE measures -- but the
README's own Calibration section says that reading holds only after
temperatures are fitted and validated, and the loader emits at load time
`RuntimeWarning: ... choice:11+=0.1006 -> 0.5. Treat confidence from the
affected entries as uncalibrated.`

The page now prints, per question shape, the temperature this checkpoint
actually applies, the bucket whose shipped value the runtime refused, and the
answers whose scaling is indistinguishable from 1.0. It also reads back the
`confidence` field the agent reported beside the page's own 1 - H/log(k)
recomputation instead of substituting the arithmetic for the field.

Measured on the shipped English checkpoint offline:

   question            type    k    bucket        applied   source
   intent              choice  6    choice:6-10   1.0000    bucket map
   is_urgent           noul    2    noul:2        1.9834    bucket map
   frustration         score   4    score:3-5     1.2514    bucket map
   refund_requested    noul    2    noul:2        1.9834    bucket map
   churn_risk          noul    2    noul:2        1.9834    bucket map
   shipped but refused by the runtime: choice:11+=0.1006 -> 0.5000
   at a temperature indistinguishable from 1.0: `intent` (1 of 5 question shapes)

   vague ticket's intent answer -- recomputed 0.1404 vs reported 0.1404,
   0.000014 apart; answer_confidence = 0.36.
Mirror the page-28 gate idiom: exec the four new helpers from the example's
own AST against a fabricated agent, with `temp_bucket`, `QTYPES` and `math`
injected from `laya.common`.

   - `test_scale_for_replays_core_lookup` re-derives `Agent.predict_batch`'s
     own `temperature_by_options.get(temp_bucket(...), temperature[QTYPES[...]])`
     on a stand-in agent and compares; the not-in-map fallback and a mapped
     temperature of exactly 1.0 are both exercised.
   - `test_option_count_matches_the_shipped_preset` proves it on real
     `laya.triage_questions()`: {intent 6, is_urgent 2, frustration 4,
     refund_requested 2, churn_risk 2}.
   - `test_clamped_buckets_names_only_a_shipped_value_the_runtime_refused`
     reproduces the loader's `choice:11+=0.1006 -> 0.5` shape and returns []
     on a self-consistent checkpoint.
   - `test_entropy_confidence_recomputes_core_exactly` asserts equality with
     `common.confidence_from_probs` on four distributions and the k<2 edge.
   - `test_page40_helpers_are_live_defer_to_core` refuses a helper that stops
     calling `temp_bucket` and hardcodes a bucket name.
   - `test_page40_drops_the_unconditional_calibration_claim` and
     `test_page40_reads_the_scaling_and_the_reported_field` are witnessed
     against an `OLD_PAGE_40` literal: each ban fires on main's wording and
     not on this page; each positive rule fails on main and passes here.

Mutation harness `laya-bench/mutate_example40_gate.py`: 7/7 caught, tree
restored True. Whole suite: `141 passed, 0 failed` via `python
tests/test_confidence.py` (the CI path).
@NandhaKishorM
NandhaKishorM merged commit 31fae85 into NandhaKishorM:main Oct 7, 2026
30 checks passed
NandhaKishorM pushed a commit that referenced this pull request Oct 7, 2026
The page measures `total_ms` and `median_ms` from `perf_counter` / `timed()` on
the same 10-case loop, prints them, and then closed with:

    Ten routing decisions cost well under a millisecond in total, and the router
    still holds zero checkpoints.

Every clause in that sentence is a quantity the page already had and did not
read back: "Ten" was hardcoded while `len(CASES)` sat right there; "well under
a millisecond" asserted a threshold that goes stale on slower hardware; "holds
zero checkpoints" asserted residency rather than reading `r.loaded`. This is
the same defect class #850 fixed on example 28's preset page and #972 fixed on
example 40's Part C, and the fix is the same shape: two pure helpers turn the
measurement into a verdict and the paragraph interpolates the verdict and the
raw number.

Measured on the current MPS run: `0.415 ms total / 10 cases`, `router.loaded
= []` after every call, so the page prints "10 routing decisions cost under
1.0 ms in total (0.415 ms measured this run), and the router holds zero
checkpoints". Lengthen `CASES` or slow the machine and the sentence
rephrases rather than lies.
@NandhaKishorM

Copy link
Copy Markdown
Owner

Merged in 0.3.29.

examples/40_caching_and_monitoring.py called answer_confidence "the calibrated probability of the answer Laya reports" in its banner, its closing paragraph and its bucket() docstring, on a checkpoint whose loader emits a RuntimeWarning at load time saying to treat confidence from the affected entries as uncalibrated. Calling it calibrated in a page about gating on it is the worst place for that claim to sit.

Reading the temperature each shape was actually scaled by off the loaded agent, and stating the calibration as conditional, is the right correction: it makes the page true for whatever checkpoint the reader has rather than for a hypothetical refitted one.

The gate pulls the four helpers out of the example's AST and execs them against a fabricated agent, with scale_for's lookup checked against core's own temp_bucket, so it drives the code the page runs instead of re-reading its prose. That is the difference between a docs test and a spell check. No weights loaded.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants