Repository navigation
fix(structured): expose answer_confidence on DecisionResult - #685
NandhaKishorM merged 2 commits into
Conversation
…omises `docs/structured.md` calls this bridge one that returns "typed values with calibrated confidence". It does not. `DecisionResult.confidence` is built from the answer's `confidence` field, and `laya/common.py` documents that field as "not calibrated: it is not what temperature scaling fits and not what the reported ECE measures. See `answer_confidence`". The two quantities are also on different scales: `tests/test_confidence.py` pins that a two-option distribution comes back as 0.90 on a `noul` and 0.53 on an equivalent `choice`, and NandhaKishorM#394 is an open issue saying a confidence threshold does not transfer across option counts, which is specifically the entropy definition's failure mode. So the number the artifact named `confidence` was never the calibrated one, and the docs' own example told users to gate on it: if result.confidence["department"] < 0.6: result.values["department"] = "human-review" Laya's own abstention gate is defined against `answer_confidence` -- NandhaKishorM#361 converged on "an opt-in threshold, off by default, that reads `answer_confidence` (`max(p)`), because that is the value the calibration figures describe and it does not drift with the number of options" -- so the structured bridge was the one surface reporting a different quantity from the gate. Worse, `confidence` defaults to `0.0` for a field that reported no confidence at all. A caller filtering this artifact to decide what to escalate cannot tell an absent confidence from a genuine zero, and escalates both for the same reason. `DecisionResult.answer_confidence` now carries the calibrated value per field, under its own name, as `None` when the answer reported no usable one. `confidence` keeps the entropy value it has always had and its default of `0.0`: NandhaKishorM#302 settled that Laya's `confidence` is defined as one minus normalised entropy and "I'll keep it rather than change its meaning for existing users", so this adds a name rather than repurposing one. A caller who needs the old field has it; a caller who filters on the calibrated one has that too, and the two cannot be confused. `laya/confidence.py` grows `calibrated_confidence(answer)`, which reads `answer_confidence` and deliberately does *not* fall back to `confidence` -- reporting the entropy number under the calibrated name would be worse than reporting nothing, because the name is what a caller filters on. The gate's own read path now calls it, so `DecisionResult` and `min_confidence` cannot disagree about what calibrated means, which is the same one-concept-one-implementation shape NandhaKishorM#533/NandhaKishorM#535 settled for the integrations' error class. `flag_low_confidence` keeps its signature, semantics and `0.0` no-op; it is public API with a test pinning that. The field is appended with a default so every existing construction of the dataclass keeps working and the first six fields keep their positions; `tests/test_structured_api.py` pins that, alongside its existing field-list pin which this updates. `docs/structured.md` now shows both numbers, states the definitions, and gates on the calibrated one. That example was the bug report. RED -> GREEN, weight-free, against 9d95567: `tests/test_structured.py` 114 passed (was 103), `tests/test_structured_api.py` 38 passed. The new assertions build a case where the two quantities rank the same two fields in opposite order, so the test fails if either number is wrong rather than merely if a field is missing. Full `ci.yml` suite list, all 62: 5 failures, all identical on unmodified 9d95567 in this environment (test_docker_entrypoint needs docker, test_env_docs is the LAYA_MAX_CONCURRENT gap, test_mcp and test_onnx_quantize need their extras, test_tokenizer_cache needs symlink privilege). Maintainer gate green: ruff, compileall, `git diff --check`, `zensical build --strict --clean`. This is independent of NandhaKishorM#679: the wrong-quantity and the `0.0`-means-absent problems both exist on 9d95567, and nothing here reads a key that NandhaKishorM#679 adds.
|
CI note: every check on this PR passes except two, and neither is this change.
I do not have admin rights on this repository, so I could not re-run that job to demonstrate it green. If it matters before review, the Green: |
`answer_confidence` is the *quantity* temperature scaling fits and every calibration figure in
this repository is computed on. It is not a claim that the number is right. The first commit
of this branch said "calibrated confidence", "calibrated max(p)", "how much should I trust
this answer", and named the helper `calibrated_confidence`. All four overstate it, and the
README is explicit that they should:
Both checkpoints are over-confident as shipped and `laya-multilingual` has no fitted
temperatures at all, so fit them before relying on these numbers ... Confidence orders
decisions; it does not establish that a decision is correct.
Reading `answer_confidence` as "about c of the answers returned at c are correct" holds only
after temperatures have been fitted and validated on held-out data for that checkpoint and
question shape. Naming a public helper after a property the shipped checkpoints do not have
reintroduces exactly the ambiguity NandhaKishorM#419 caught: the quantity that calibration targets is not
the same claim as a calibrated model.
What changes:
- `calibrated_confidence(answer)` -> `answer_confidence_value(answer)`. A name that says which
field it reads and asserts nothing about it. Still not in the top-level `__all__`.
- Every docstring and comment that said "calibrated" now says what the number is (`max(p)`, the
probability mass on the reported answer; the quantity temperature scaling fits and ECE
evaluates) and states the condition under which it means c-approximates-correctness. The
`DecisionResult` docstring, `laya/confidence.py`, `docs/structured.md` and the tests all say
the same thing, including that the shipped checkpoints are over-confident.
- The product argument does not change and is now stated more precisely. `answer_confidence` is
the number to filter on **because it is the same quantity `min_confidence` compares against
and the eval and calibration stack measures** -- not because it is more trustworthy. That is a
stronger claim than the old one, because it is checkable: the gate and this field read one
helper, so they cannot disagree about which number is being reported.
- `docs/structured.md` no longer claims the bridge returns "calibrated confidence", and its
example says explicitly that filtering on the right number is not the same as it being a
trustworthy probability, with links to the README's Calibration and Honest limits sections.
- Fixed a dangling `[Calibration](#calibration)` anchor the docs gate caught; that section is in
the README, so it is a cross-doc link now.
The field, the `None` semantics, the no-fallback rule and the `confidence` field's unchanged
value and default all stay as they were -- this is a naming and semantics correction only.
`flag_low_confidence` still has the same signature, semantics and 0.0 no-op.
Unchanged: tests/test_structured.py 114 passed, tests/test_structured_api.py 38,
tests/test_confidence.py 59, tests/test_packaging.py 106, tests/test_structured_docs.py 87,
tests/test_doc_tables.py 109, tests/test_hooks_api.py 387, tests/test_hooks.py 204,
tests/test_langchain.py 207. Full ci.yml suite list, 62 suites, 5 failures -- the same five
that fail identically on unmodified 9d95567 in this environment. ruff, compileall,
git diff --check and zensical build --strict --clean all clean.
|
Amendment (
What changed:
Unchanged: the field, the Re-verified after the amendment: |
|
Passed checks and merged. Thank you @Bruce-Yii! It ships in 0.3.22. |
…te ran The gate is a policy, and a policy whose application cannot be observed is not one. `low_confidence` is written only when the gate fires, so its absence cannot tell a caller that a gate ran and this answer cleared it apart from no gate running at all. That leaves three things unanswerable from a log: what fraction of decisions abstained, whether the gate was in effect, and which threshold produced the batch, since `flag_low_confidence` consumes the threshold and drops it. With `min_confidence` set, every answer now also carries `abstention` -- `passed`, `abstained`, or `unevaluated` -- plus `abstention_threshold` echoing the threshold. `unevaluated` is the case a boolean cannot express: the gate ran and the answer carried no usable confidence, so it could not decide. Reporting that as a pass would be the same lie as reporting it as a flag. With `min_confidence` unset, nothing is written at all -- no `abstention`, no threshold, no flag. An ungated call returns exactly the payload it returned before, so the response schema is unchanged for the callers who never asked to be gated, and the presence of `abstention` is what tells the two cases apart. That is why the vocabulary has no `not_configured` member: a sentinel for the unconfigured case would put that information back into the payload for every caller, which is the thing this avoids. The callers pass the threshold unconditionally and the branch lives in the one place that owns the contract, rather than at each of the ten call sites where an `if min_confidence is not None:` guard would leave a path silently reporting nothing. The flag stays `flag_low_confidence`'s; this delegates rather than re-implementing the rule, so the boolean and the reported state cannot drift apart. Also fixes a leak in the helper both share. It returned whatever `confidence` held when the validity test failed, so an answer with no `answer_confidence` and a NaN `confidence` returned NaN. `flag_low_confidence` never noticed -- `NaN < x` is False and `None is not None` is also False, so it flagged either way -- but a caller that reads the number cannot tell "nothing to gate on" from "a gate ran on a NaN", and those two now report differently. Only the private helper's return value changes; no public behaviour does. Conflict with merged NandhaKishorM#685 resolved by adopting its `answer_confidence_value` and dropping this branch's duplicate `_gate_confidence`, so there is one helper rather than two that can disagree. NandhaKishorM#685's semantics are preserved: `answer_confidence` stays the per-field quantity, and nothing here calls it calibrated. Closes NandhaKishorM#361
…te ran The gate is a policy, and a policy whose application cannot be observed is not one. `low_confidence` is written only when the gate fires, so its absence cannot tell a caller that a gate ran and this answer cleared it apart from no gate running at all. That leaves three things unanswerable from a log: what fraction of decisions abstained, whether the gate was in effect, and which threshold produced the batch, since `flag_low_confidence` consumes the threshold and drops it. With `min_confidence` set, every answer now also carries `abstention` -- `passed`, `abstained`, or `unevaluated` -- plus `abstention_threshold` echoing the threshold. `unevaluated` is the case a boolean cannot express: the gate ran and the answer carried no usable confidence, so it could not decide. Reporting that as a pass would be the same lie as reporting it as a flag. With `min_confidence` unset, nothing is written at all -- no `abstention`, no threshold, no flag. An ungated call returns exactly the payload it returned before, so the response schema is unchanged for the callers who never asked to be gated, and the presence of `abstention` is what tells the two cases apart. That is why the vocabulary has no `not_configured` member: a sentinel for the unconfigured case would put that information back into the payload for every caller, which is the thing this avoids. The callers pass the threshold unconditionally and the branch lives in the one place that owns the contract, rather than at each of the ten call sites where an `if min_confidence is not None:` guard would leave a path silently reporting nothing. The flag stays `flag_low_confidence`'s; this delegates rather than re-implementing the rule, so the boolean and the reported state cannot drift apart. Also fixes a leak in the helper both share. It returned whatever `confidence` held when the validity test failed, so an answer with no `answer_confidence` and a NaN `confidence` returned NaN. `flag_low_confidence` never noticed -- `NaN < x` is False and `None is not None` is also False, so it flagged either way -- but a caller that reads the number cannot tell "nothing to gate on" from "a gate ran on a NaN", and those two now report differently. Only the private helper's return value changes; no public behaviour does. Conflict with merged NandhaKishorM#685 resolved by adopting its `answer_confidence_value` and dropping this branch's duplicate `_gate_confidence`, so there is one helper rather than two that can disagree. NandhaKishorM#685's semantics are preserved: `answer_confidence` stays the per-field quantity, and nothing here calls it calibrated. Closes NandhaKishorM#361
- Add decideBatch and snake_case alias decide_batch to laya-ts (structured decisions over batched states).
- Follow both Agent convention (states, questions, opts) and Router convention (requests mapped with { state, questions } to predictBatch).
- Support minConfidence / min_confidence abstention gating on batched states, projecting low-confidence answers to null.
- Expose answer_confidence (calibrated max(p)) on DecisionResult, matching Python parity from PR NandhaKishorM#685.
- Expose decideBatch and decide_batch on Agent and Router.
- Add test coverage across schema projection, explicit questions, details, confidence gating, and error handling.
- Add decideBatch to laya-ts (structured decisions over batched states).
- Follow both Agent convention (states, questions, opts) and Router convention (requests mapped with { state, questions } to predictBatch).
- Support minConfidence abstention gating on batched states, projecting low-confidence answers to null.
- Expose answer_confidence (calibrated max(p)) on DecisionResult, matching Python parity from PR NandhaKishorM#685.
- Expose decideBatch on Agent and Router.
- Add test coverage across schema projection, explicit questions, details, confidence gating, and error handling.
- Add decideBatch to laya-ts (structured decisions over batched states).
- Follow both Agent convention (states, questions, opts) and Router convention (requests mapped with { state, questions } to predictBatch).
- Support minConfidence abstention gating on batched states, projecting low-confidence answers to null.
- Expose answer_confidence (calibrated max(p)) on DecisionResult, matching Python PR NandhaKishorM#685 parity.
- Expose decideBatch on Agent and Router.
- Add test coverage across schema projection, explicit questions, details, confidence gating, and error handling.
The user decision
You run
laya.decide(..., return_details=True), get aDecisionResultper state, and filter it todecide what routes automatically and what goes to a human. That filter is the decision. Here is the
shipped example from
docs/structured.md, verbatim:It gates on
result.confidence.The failure mode
DecisionResult.confidenceis a different quantity. It is built from the answer'sconfidencefield, andlaya/common.py:confidence_from_probsdocuments that field as:The two are also on different scales, which
tests/test_confidence.pypins: a two-optiondistribution comes back as 0.90 on a
nouland 0.53 on an equivalentchoice, and the samedocumented threshold separates them at 0.85. Issue #394 is open on exactly this — "a confidence
threshold does not transfer across option counts" — which is the entropy definition's specific
failure mode, because
log(k)is in the denominator.So the number the example filters on moves with the shape of the question, not with how right the
answer is. Add a fourth option to
departmentand every field's reported confidence shifts, forreasons that have nothing to do with the decision.
And
confidencedefaults to0.0.structured.pyreadsanswer.get("confidence", 0.0), so afield that reported no confidence at all, a field the model did not return, and a field whose
confidence really is zero are one value. A caller filtering "below 0.6, escalate" escalates all
three identically and cannot say which it was looking at.
Meanwhile Laya's own abstention gate is defined against the other number. On #361 the maintainer
converged on:
So the structured bridge is the one surface reporting a different quantity from the gate that
min_confidenceimplements, while the docs pointed the reader at the other one.Maintainer precedent
The closest precedent is #621, where he was asked to change what an existing property reported and
declined — while naming the shape he would take:
And #302, which settles that an existing field's meaning is not available to be repurposed:
This PR is exactly that shape: additive, read-only, and
confidencekeeps its value and its default.Also #533/#535, where three integrations were given one shared error class so they could not
drift — here
DecisionResultand the gate read the reported quantity through one helper.Boundary inference, with the falsifier. High confidence that this fits: it adds no policy, chooses
no threshold, owns no queue, executes nothing, and leaves every application-owned field alone. The
falsifier is
docs/staged-adoption.md:16-18, which lists what a shadow record should contain — andevery item on it (the incumbent action, the reviewed outcome, the review decision) is something Laya
does not know. This changes only the half Laya does know, in an artifact Laya already returns, so it
is not a claim on the record. If the maintainer reads
DecisionResultas part of the application-ownedhandoff, this should have been a docs correction only, and I would rather be told that than be wrong.
On the name.
answer_confidenceis the quantity temperature scaling fits and everycalibration figure here is computed on. It is not a claim that the number is right: reading it
as "about c of the answers returned at c are correct" holds only after temperatures are
fitted and validated on held-out data for that checkpoint and question shape, and the shipped
checkpoints are over-confident as shipped. So the reason to report it is not that it is more
trustworthy -- it is that it is the same number the gate and the eval harness use, which is
checkable, rather than a property that has to be taken on trust.
The Laya contract
DecisionResultreportsmax(p)under its own name, and a field that reported nousable confidence is
None— a value that is not0.0, because "nobody told me" and "they told mezero" are different facts and a caller acting on them should not have to guess which one it is.
DecisionResult.answer_confidence— themax(p)per field,Nonewhen absent/unusable.DecisionResult.confidence— unchanged: the entropy value, unchanged default of0.0.laya.confidence.answer_confidence_value, which deliberately does not fall back toconfidence. Reporting the entropy number under this name would be worse than reportingnothing, because the name is what the caller filters on. The gate's own read path now calls the same
helper, so the artifact and
min_confidencecannot disagree.Implementation
laya/confidence.py—answer_confidence_value(answer); the gate's private_gate_confidencenowdelegates to it and keeps its own entropy fallback.
flag_low_confidenceis untouched: samesignature, same semantics, same
0.0no-op, still public API with a test pinning that.laya/structured.py— theanswer_confidencefield, populated in_details.working and the first six fields keep their positions.
tests/test_structured_api.pypins thatalongside its field-list check.
docs/structured.md— shows both numbers, states both definitions, and gates on the one the gate uses.That example was the bug report, so fixing it is part of the fix.
Validation
RED → GREEN, weight-free, against
9d95567:The new assertions do not merely check that a field exists. They build a case where the two
quantities rank the same two fields in opposite order, so the test fails if either number is
wrong, and they pin that absent /
True/NaNall land onNonerather than0.0.Full
ci.ymlsuite list — all 62 suites, extracted from the workflow rather than hand-picked:Maintainer gate: ruff,
compileall,git diff --check,zensical build --strict --clean— all clean.Independent of #664 and #679. Both the wrong-quantity problem and the
0.0-means-absent problemexist on
9d95567, and nothing here reads a key that either of those PRs adds. This can merge in anyorder.
Collision
No open PR touches
DecisionResultor the confidence fields. Two open PRs do touch the samefiles, though, so this is worth stating rather than glossing: #683 and #666 both change
laya/structured.pyandtests/test_structured.py(andtests/test_structured_docs.py, whichthis one does not touch). They are about the score-minimum projection in
_project; this isabout
DecisionResultand_details. Different functions, so a merge should be clean ineither order, but the files overlap. #657, #644, #667, #476 and #419 touch none of these files.
Non-goals
confidenceis not changed, redefined, or deprecated. It has a real use — "how concentrated isthis distribution" — and /v1/systemone: 128-option choices get 422, a null score level comes back as a null legend value, and confidence differs from TypeSafe's formula #302 settled its meaning. Two names, two definitions, both reported.
min_confidencestill reads the answer's ownanswer_confidence; nothing here makes Laya choose one.valuesprojection. An abstained field is stillNoneinvalues; changingthe schema-shaped output would break its contract.
decide's signature,laya-ts, or any surface other thanDecisionResult. feat(ts): port confidence-based abstention gating on answer_confidence (#361) #644and fix(ts): honour score minimum in structured projection (#663) #676 are working in
laya-ts/src/structured.ts; there is a parity question there, and I wouldrather raise it in review than bundle a second language port into a contract fix.
Tradeoffs
A second number in the artifact. Two confidences where there was one is more to understand. I
think it is right — they are different quantities with different scales and the existing one has a
documented use — but a maintainer who wants a single field would reasonably disagree.
Noneis new in aDict[str, float]. Callers iteratinganswer_confidencemust handle a missingentry as
Nonerather than assume a float. That is the point, and it is also a new thing todocument.
answer_confidence_valueis not exported fromlaya. It is public inlaya.confidenceand used bythe library; I kept it off the top-level
__all__becausetest_packaging.pyrequires every exportedname to be documented, and an exported helper nobody needs to call is API surface for nothing. Say the
word if you would rather it were top-level.
What would falsify this
DecisionResultis application-owned in the maintainer's view, this should have been adocumentation correction and the field is scope creep.
docs/structured.mdchange alone and I would withdraw the field.DecisionResulton confidence in practice, the field is a nicety. The argumentis that it is the documented way to decide what to automate, and it currently points at the wrong
number — but I have not measured how many projects do this, and I have no evidence that they do.