Skip to content

fix(structured): expose answer_confidence on DecisionResult - #685

Merged
NandhaKishorM merged 2 commits into
NandhaKishorM:mainfrom
Bruce-Yii:feat/decision-calibrated-confidence
Sep 29, 2026
Merged

NandhaKishorM merged 2 commits into
NandhaKishorM:mainfrom
Bruce-Yii:feat/decision-calibrated-confidence

Conversation

@Bruce-Yii

@Bruce-Yii Bruce-Yii commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

The user decision

Which of my structured decisions are safe to automate?

You run laya.decide(..., return_details=True), get a DecisionResult per state, and filter it to
decide what routes automatically and what goes to a human. That filter is the decision. Here is the
shipped example from docs/structured.md, verbatim:

result = agent.decide(state, schema=Ticket, return_details=True)
if result.confidence["department"] < 0.6:
    result.values["department"] = "human-review"

It gates on result.confidence.

The failure mode

DecisionResult.confidence is a different quantity. It is built from the answer's confidence field, and
laya/common.py:confidence_from_probs documents that field as:

"Normalized Shannon entropy confidence: 1 - H(p) / log(k). How concentrated the whole distribution
is. Useful, but not calibrated: it is not what temperature scaling fits and not what the reported
ECE measures.
See answer_confidence."

The two are also on different scales, which tests/test_confidence.py pins: a two-option
distribution comes back as 0.90 on a noul and 0.53 on an equivalent choice, and the same
documented threshold separates them at 0.85. Issue #394 is open on exactly this — "a confidence
threshold does not transfer across option counts" — which is the entropy definition's specific
failure mode, because log(k) is in the denominator.

So the number the example filters on moves with the shape of the question, not with how right the
answer is. Add a fourth option to department and every field's reported confidence shifts, for
reasons that have nothing to do with the decision.

And confidence defaults to 0.0. structured.py reads answer.get("confidence", 0.0), so a
field that reported no confidence at all, a field the model did not return, and a field whose
confidence really is zero are one value. A caller filtering "below 0.6, escalate" escalates all
three identically and cannot say which it was looking at.

Meanwhile Laya's own abstention gate is defined against the other number. On #361 the maintainer
converged on:

"an opt-in threshold, off by default, that reads answer_confidence (max(p)), because that is
the value the calibration figures describe and it does not drift with the number of options
"

So the structured bridge is the one surface reporting a different quantity from the gate that
min_confidence implements, while the docs pointed the reader at the other one.

Maintainer precedent

The closest precedent is #621, where he was asked to change what an existing property reported and
declined — while naming the shape he would take:

"Changing what dtype returns would affect the autocast call itself, so I'd rather not do that.
I'd take a PR that either documents dtype as the autocast target … or adds a small read-only
way to ask
what precision a call with N rows runs in."

And #302, which settles that an existing field's meaning is not available to be repurposed:

"Laya's confidence is defined as 1 minus normalised entropy, and the thresholds in Laya's docs are
calibrated to that definition, so I'll keep it rather than change its meaning for existing users."

This PR is exactly that shape: additive, read-only, and confidence keeps its value and its default.
Also #533/#535, where three integrations were given one shared error class so they could not
drift — here DecisionResult and the gate read the reported quantity through one helper.

Boundary inference, with the falsifier. High confidence that this fits: it adds no policy, chooses
no threshold, owns no queue, executes nothing, and leaves every application-owned field alone. The
falsifier is docs/staged-adoption.md:16-18, which lists what a shadow record should contain — and
every item on it (the incumbent action, the reviewed outcome, the review decision) is something Laya
does not know. This changes only the half Laya does know, in an artifact Laya already returns, so it
is not a claim on the record. If the maintainer reads DecisionResult as part of the application-owned
handoff, this should have been a docs correction only, and I would rather be told that than be wrong.

On the name. answer_confidence is the quantity temperature scaling fits and every
calibration figure here is computed on. It is not a claim that the number is right: reading it
as "about c of the answers returned at c are correct" holds only after temperatures are
fitted and validated on held-out data for that checkpoint and question shape, and the shipped
checkpoints are over-confident as shipped. So the reason to report it is not that it is more
trustworthy -- it is that it is the same number the gate and the eval harness use, which is
checkable, rather than a property that has to be taken on trust.

The Laya contract

DecisionResult reports max(p) under its own name, and a field that reported no
usable confidence is None — a value that is not 0.0, because "nobody told me" and "they told me
zero" are different facts and a caller acting on them should not have to guess which one it is.

  • DecisionResult.answer_confidence — the max(p) per field, None when absent/unusable.
  • DecisionResult.confidence — unchanged: the entropy value, unchanged default of 0.0.
  • Read through laya.confidence.answer_confidence_value, which deliberately does not fall back to
    confidence. Reporting the entropy number under this name would be worse than reporting
    nothing, because the name is what the caller filters on. The gate's own read path now calls the same
    helper, so the artifact and min_confidence cannot disagree.

Implementation

  • laya/confidence.py — answer_confidence_value(answer); the gate's private _gate_confidence now
    delegates to it and keeps its own entropy fallback. flag_low_confidence is untouched: same
    signature, same semantics, same 0.0 no-op, still public API with a test pinning that.
  • laya/structured.py — the answer_confidence field, populated in _details.
  • The field is appended with a default, so every existing construction of the dataclass keeps
    working and the first six fields keep their positions. tests/test_structured_api.py pins that
    alongside its field-list check.
  • docs/structured.md — shows both numbers, states both definitions, and gates on the one the gate uses.
    That example was the bug report, so fixing it is part of the fix.

Validation

RED → GREEN, weight-free, against 9d95567:

tests/test_structured.py       114 passed, 0 failed   (was 103)
tests/test_structured_api.py    38 passed, 0 failed

The new assertions do not merely check that a field exists. They build a case where the two
quantities rank the same two fields in opposite order, so the test fails if either number is
wrong, and they pin that absent / True / NaN all land on None rather than 0.0.

Full ci.yml suite list — all 62 suites, extracted from the workflow rather than hand-picked:

62 suites, 5 failures -- all five fail identically on unmodified 9d95567 in this
environment (verified by stashing and re-running each):
  test_docker_entrypoint.py  needs docker
  test_env_docs.py           the LAYA_MAX_CONCURRENT doc gap
  test_mcp.py                needs the mcp extra
  test_onnx_quantize.py      needs the onnx extra
  test_tokenizer_cache.py    needs symlink privilege (WinError 1314)

Maintainer gate: ruff, compileall, git diff --check, zensical build --strict --clean — all clean.

Independent of #664 and #679. Both the wrong-quantity problem and the 0.0-means-absent problem
exist on 9d95567, and nothing here reads a key that either of those PRs adds. This can merge in any
order.

Collision

No open PR touches DecisionResult or the confidence fields. Two open PRs do touch the same
files, though, so this is worth stating rather than glossing: #683 and #666 both change
laya/structured.py and tests/test_structured.py (and tests/test_structured_docs.py, which
this one does not touch). They are about the score-minimum projection in _project; this is
about DecisionResult and _details. Different functions, so a merge should be clean in
either order, but the files overlap. #657, #644, #667, #476 and #419 touch none of these files.

Non-goals

Tradeoffs

A second number in the artifact. Two confidences where there was one is more to understand. I
think it is right — they are different quantities with different scales and the existing one has a
documented use — but a maintainer who wants a single field would reasonably disagree.

None is new in a Dict[str, float]. Callers iterating answer_confidence must handle a missing
entry as None rather than assume a float. That is the point, and it is also a new thing to
document.

answer_confidence_value is not exported from laya. It is public in laya.confidence and used by
the library; I kept it off the top-level __all__ because test_packaging.py requires every exported
name to be documented, and an exported helper nobody needs to call is API surface for nothing. Say the
word if you would rather it were top-level.

What would falsify this

  • If DecisionResult is application-owned in the maintainer's view, this should have been a
    documentation correction and the field is scope creep.
  • If reporting both numbers is considered worse than fixing the one example, then the fix is the
    docs/structured.md change alone and I would withdraw the field.
  • If nobody filters a DecisionResult on confidence in practice, the field is a nicety. The argument
    is that it is the documented way to decide what to automate, and it currently points at the wrong
    number — but I have not measured how many projects do this, and I have no evidence that they do.

…omises

`docs/structured.md` calls this bridge one that returns "typed values with calibrated
confidence". It does not. `DecisionResult.confidence` is built from the answer's `confidence`
field, and `laya/common.py` documents that field as "not calibrated: it is not what temperature
scaling fits and not what the reported ECE measures. See `answer_confidence`". The two
quantities are also on different scales: `tests/test_confidence.py` pins that a two-option
distribution comes back as 0.90 on a `noul` and 0.53 on an equivalent `choice`, and NandhaKishorM#394 is an
open issue saying a confidence threshold does not transfer across option counts, which is
specifically the entropy definition's failure mode.

So the number the artifact named `confidence` was never the calibrated one, and the docs' own
example told users to gate on it:

    if result.confidence["department"] < 0.6:
        result.values["department"] = "human-review"

Laya's own abstention gate is defined against `answer_confidence` -- NandhaKishorM#361 converged on "an
opt-in threshold, off by default, that reads `answer_confidence` (`max(p)`), because that is
the value the calibration figures describe and it does not drift with the number of options" --
so the structured bridge was the one surface reporting a different quantity from the gate.

Worse, `confidence` defaults to `0.0` for a field that reported no confidence at all. A caller
filtering this artifact to decide what to escalate cannot tell an absent confidence from a
genuine zero, and escalates both for the same reason.

`DecisionResult.answer_confidence` now carries the calibrated value per field, under its own
name, as `None` when the answer reported no usable one. `confidence` keeps the entropy value it
has always had and its default of `0.0`: NandhaKishorM#302 settled that Laya's `confidence` is defined as
one minus normalised entropy and "I'll keep it rather than change its meaning for existing
users", so this adds a name rather than repurposing one. A caller who needs the old field has
it; a caller who filters on the calibrated one has that too, and the two cannot be confused.

`laya/confidence.py` grows `calibrated_confidence(answer)`, which reads `answer_confidence` and
deliberately does *not* fall back to `confidence` -- reporting the entropy number under the
calibrated name would be worse than reporting nothing, because the name is what a caller filters
on. The gate's own read path now calls it, so `DecisionResult` and `min_confidence` cannot
disagree about what calibrated means, which is the same one-concept-one-implementation shape
NandhaKishorM#533/NandhaKishorM#535 settled for the integrations' error class. `flag_low_confidence` keeps its signature,
semantics and `0.0` no-op; it is public API with a test pinning that.

The field is appended with a default so every existing construction of the dataclass keeps
working and the first six fields keep their positions; `tests/test_structured_api.py` pins
that, alongside its existing field-list pin which this updates.

`docs/structured.md` now shows both numbers, states the definitions, and gates on the
calibrated one. That example was the bug report.

RED -> GREEN, weight-free, against 9d95567: `tests/test_structured.py` 114 passed (was 103),
`tests/test_structured_api.py` 38 passed. The new assertions build a case where the two
quantities rank the same two fields in opposite order, so the test fails if either number is
wrong rather than merely if a field is missing. Full `ci.yml` suite list, all 62: 5 failures,
all identical on unmodified 9d95567 in this environment (test_docker_entrypoint needs docker,
test_env_docs is the LAYA_MAX_CONCURRENT gap, test_mcp and test_onnx_quantize need their extras,
test_tokenizer_cache needs symlink privilege). Maintainer gate green: ruff, compileall,
`git diff --check`, `zensical build --strict --clean`.

This is independent of NandhaKishorM#679: the wrong-quantity and the `0.0`-means-absent problems both exist on
9d95567, and nothing here reads a key that NandhaKishorM#679 adds.
@Bruce-Yii

Copy link
Copy Markdown
Contributor Author

CI note: every check on this PR passes except two, and neither is this change.

dependency CVEs — the repo-wide pip-audit --strict audit. It cannot resolve the PyTorch CPU index's 2.14.0+cpu, a local version absent from PyPI. It fails identically on #664, #679, #661, #653, and on unrelated PRs that touch nothing.

smoke (ubuntu-latest) — a PyPI network timeout, not a dependency conflict. The log is five ReadTimeoutErrors reaching pypi.org while resolving click, then ResolutionImpossible because the index was never reachable:

WARNING: Retrying (total=4 ...) after connection broken by 'ReadTimeoutError("HTTPSConnectionPool(host='"'"'pypi.org'"'"', port=443): Read timed out. (read timeout=15)")': /simple/click/
...
ERROR: Cannot install laya because these package versions have conflicting dependencies.
ERROR: ResolutionImpossible: for help visit https://pip.pypa.io/en/latest/topics/dependency-resolution/

smoke (ubuntu-24.04-arm) and spark-build passed in the same run and build the same image, and tests passed on py3.10 / 3.11 / 3.12 / 3.13 / windows. This diff touches no dependency or build file — docs/structured.md, laya/confidence.py, laya/structured.py, tests/test_structured.py, tests/test_structured_api.py — so there is no dependency change for it to conflict with.

I do not have admin rights on this repository, so I could not re-run that job to demonstrate it green. If it matters before review, the smoke workflow can be re-triggered on demand.

Green: tests (5 matrices), evals, docs (strict build), lint, CodeQL, onnx export, smoke (ubuntu-24.04-arm), spark-build, http, package, Chinese benchmark audit, no unsafe deserialization, secret scan.

`answer_confidence` is the *quantity* temperature scaling fits and every calibration figure in
this repository is computed on. It is not a claim that the number is right. The first commit
of this branch said "calibrated confidence", "calibrated max(p)", "how much should I trust
this answer", and named the helper `calibrated_confidence`. All four overstate it, and the
README is explicit that they should:

    Both checkpoints are over-confident as shipped and `laya-multilingual` has no fitted
    temperatures at all, so fit them before relying on these numbers ... Confidence orders
    decisions; it does not establish that a decision is correct.

Reading `answer_confidence` as "about c of the answers returned at c are correct" holds only
after temperatures have been fitted and validated on held-out data for that checkpoint and
question shape. Naming a public helper after a property the shipped checkpoints do not have
reintroduces exactly the ambiguity NandhaKishorM#419 caught: the quantity that calibration targets is not
the same claim as a calibrated model.

What changes:

- `calibrated_confidence(answer)` -> `answer_confidence_value(answer)`. A name that says which
  field it reads and asserts nothing about it. Still not in the top-level `__all__`.
- Every docstring and comment that said "calibrated" now says what the number is (`max(p)`, the
  probability mass on the reported answer; the quantity temperature scaling fits and ECE
  evaluates) and states the condition under which it means c-approximates-correctness. The
  `DecisionResult` docstring, `laya/confidence.py`, `docs/structured.md` and the tests all say
  the same thing, including that the shipped checkpoints are over-confident.
- The product argument does not change and is now stated more precisely. `answer_confidence` is
  the number to filter on **because it is the same quantity `min_confidence` compares against
  and the eval and calibration stack measures** -- not because it is more trustworthy. That is a
  stronger claim than the old one, because it is checkable: the gate and this field read one
  helper, so they cannot disagree about which number is being reported.
- `docs/structured.md` no longer claims the bridge returns "calibrated confidence", and its
  example says explicitly that filtering on the right number is not the same as it being a
  trustworthy probability, with links to the README's Calibration and Honest limits sections.
- Fixed a dangling `[Calibration](#calibration)` anchor the docs gate caught; that section is in
  the README, so it is a cross-doc link now.

The field, the `None` semantics, the no-fallback rule and the `confidence` field's unchanged
value and default all stay as they were -- this is a naming and semantics correction only.
`flag_low_confidence` still has the same signature, semantics and 0.0 no-op.

Unchanged: tests/test_structured.py 114 passed, tests/test_structured_api.py 38,
tests/test_confidence.py 59, tests/test_packaging.py 106, tests/test_structured_docs.py 87,
tests/test_doc_tables.py 109, tests/test_hooks_api.py 387, tests/test_hooks.py 204,
tests/test_langchain.py 207. Full ci.yml suite list, 62 suites, 5 failures -- the same five
that fail identically on unmodified 9d95567 in this environment. ruff, compileall,
git diff --check and zensical build --strict --clean all clean.
@Bruce-Yii

Copy link
Copy Markdown
Contributor Author

Amendment (acb2bef): the first commit of this branch said "calibrated confidence", "calibrated max(p)", "how much should I trust this answer", and named the helper calibrated_confidence. All four overstate what the number is, and the README is explicit that they should:

Both checkpoints are over-confident as shipped and laya-multilingual has no fitted temperatures at all, so fit them before relying on these numbers ... Confidence orders decisions; it does not establish that a decision is correct.

answer_confidence is the quantity temperature scaling fits and every calibration figure in this repository is computed on. Reading it as "about c of the answers returned at c are correct" holds only after temperatures have been fitted and validated on held-out data for that checkpoint and question shape. Naming a public helper after a property the shipped checkpoints do not have reintroduces exactly the ambiguity #419 caught — the quantity calibration targets is not the same claim as a calibrated model.

What changed:

  • calibrated_confidence(answer) → answer_confidence_value(answer). A name that says which field it reads and asserts nothing about it. Still not in the top-level __all__.
  • Every docstring and comment that said "calibrated" now says what the number is (max(p), the probability mass on the reported answer; the quantity temperature scaling fits and ECE evaluates) and states the condition under which it means c≈correctness, including that the shipped checkpoints are over-confident. DecisionResult's docstring, laya/confidence.py, docs/structured.md and the tests all say the same thing.
  • The product argument is now stated more precisely, and better: answer_confidence is the number to filter on because it is the same quantity min_confidence compares against and the eval/calibration stack measures — a checkable claim, since the gate and this field read one helper — rather than because it is more trustworthy.
  • docs/structured.md no longer implies a calibration guarantee, and its example says explicitly that filtering on the right number is not the same as it being a trustworthy probability, linking the README's Calibration and Honest limits sections.
  • Fixed a dangling [Calibration](#calibration) anchor that tests/test_packaging.py caught — that section lives in the README, so it is a cross-doc link now.

Unchanged: the field, the None semantics, the no-fallback rule, and confidence's value and default. flag_low_confidence still has the same signature, semantics and 0.0 no-op.

Re-verified after the amendment: test_structured.py 114, test_structured_api.py 38, test_confidence.py 59, test_packaging.py 106, test_structured_docs.py 87, test_doc_tables.py 109, test_hooks_api.py 387, test_hooks.py 204, test_langchain.py 207. Full ci.yml list, 62 suites, the same 5 failures that fail identically on unmodified 9d95567 here. ruff, compileall, git diff --check and zensical build --strict --clean clean.

@Bruce-Yii Bruce-Yii changed the title fix(structured): report the calibrated confidence \DecisionResult\ promises fix(structured): expose answer_confidence on DecisionResult Sep 28, 2026
@NandhaKishorM
NandhaKishorM merged commit 0f55bf6 into NandhaKishorM:main Sep 29, 2026
22 of 23 checks passed
@NandhaKishorM

Copy link
Copy Markdown
Owner

Passed checks and merged. Thank you @Bruce-Yii! It ships in 0.3.22.

Bruce-Yii added a commit to Bruce-Yii/laya that referenced this pull request Sep 30, 2026
…te ran

The gate is a policy, and a policy whose application cannot be observed is not
one. `low_confidence` is written only when the gate fires, so its absence
cannot tell a caller that a gate ran and this answer cleared it apart from no
gate running at all. That leaves three things unanswerable from a log: what
fraction of decisions abstained, whether the gate was in effect, and which
threshold produced the batch, since `flag_low_confidence` consumes the
threshold and drops it.

With `min_confidence` set, every answer now also carries `abstention` --
`passed`, `abstained`, or `unevaluated` -- plus `abstention_threshold` echoing
the threshold. `unevaluated` is the case a boolean cannot express: the gate ran
and the answer carried no usable confidence, so it could not decide. Reporting
that as a pass would be the same lie as reporting it as a flag.

With `min_confidence` unset, nothing is written at all -- no `abstention`, no
threshold, no flag. An ungated call returns exactly the payload it returned
before, so the response schema is unchanged for the callers who never asked to
be gated, and the presence of `abstention` is what tells the two cases apart.
That is why the vocabulary has no `not_configured` member: a sentinel for the
unconfigured case would put that information back into the payload for every
caller, which is the thing this avoids. The callers pass the threshold
unconditionally and the branch lives in the one place that owns the contract,
rather than at each of the ten call sites where an `if min_confidence is not
None:` guard would leave a path silently reporting nothing.

The flag stays `flag_low_confidence`'s; this delegates rather than
re-implementing the rule, so the boolean and the reported state cannot drift
apart.

Also fixes a leak in the helper both share. It returned whatever `confidence`
held when the validity test failed, so an answer with no `answer_confidence`
and a NaN `confidence` returned NaN. `flag_low_confidence` never noticed --
`NaN < x` is False and `None is not None` is also False, so it flagged either
way -- but a caller that reads the number cannot tell "nothing to gate on" from
"a gate ran on a NaN", and those two now report differently. Only the private
helper's return value changes; no public behaviour does.

Conflict with merged NandhaKishorM#685 resolved by adopting its `answer_confidence_value`
and dropping this branch's duplicate `_gate_confidence`, so there is one
helper rather than two that can disagree. NandhaKishorM#685's semantics are preserved:
`answer_confidence` stays the per-field quantity, and nothing here calls it
calibrated.

Closes NandhaKishorM#361
Bruce-Yii added a commit to Bruce-Yii/laya that referenced this pull request Sep 30, 2026
…te ran

The gate is a policy, and a policy whose application cannot be observed is not
one. `low_confidence` is written only when the gate fires, so its absence
cannot tell a caller that a gate ran and this answer cleared it apart from no
gate running at all. That leaves three things unanswerable from a log: what
fraction of decisions abstained, whether the gate was in effect, and which
threshold produced the batch, since `flag_low_confidence` consumes the
threshold and drops it.

With `min_confidence` set, every answer now also carries `abstention` --
`passed`, `abstained`, or `unevaluated` -- plus `abstention_threshold` echoing
the threshold. `unevaluated` is the case a boolean cannot express: the gate ran
and the answer carried no usable confidence, so it could not decide. Reporting
that as a pass would be the same lie as reporting it as a flag.

With `min_confidence` unset, nothing is written at all -- no `abstention`, no
threshold, no flag. An ungated call returns exactly the payload it returned
before, so the response schema is unchanged for the callers who never asked to
be gated, and the presence of `abstention` is what tells the two cases apart.
That is why the vocabulary has no `not_configured` member: a sentinel for the
unconfigured case would put that information back into the payload for every
caller, which is the thing this avoids. The callers pass the threshold
unconditionally and the branch lives in the one place that owns the contract,
rather than at each of the ten call sites where an `if min_confidence is not
None:` guard would leave a path silently reporting nothing.

The flag stays `flag_low_confidence`'s; this delegates rather than
re-implementing the rule, so the boolean and the reported state cannot drift
apart.

Also fixes a leak in the helper both share. It returned whatever `confidence`
held when the validity test failed, so an answer with no `answer_confidence`
and a NaN `confidence` returned NaN. `flag_low_confidence` never noticed --
`NaN < x` is False and `None is not None` is also False, so it flagged either
way -- but a caller that reads the number cannot tell "nothing to gate on" from
"a gate ran on a NaN", and those two now report differently. Only the private
helper's return value changes; no public behaviour does.

Conflict with merged NandhaKishorM#685 resolved by adopting its `answer_confidence_value`
and dropping this branch's duplicate `_gate_confidence`, so there is one
helper rather than two that can disagree. NandhaKishorM#685's semantics are preserved:
`answer_confidence` stays the per-field quantity, and nothing here calls it
calibrated.

Closes NandhaKishorM#361
Angboo pushed a commit to Angboo/laya that referenced this pull request Sep 30, 2026
- Add decideBatch and snake_case alias decide_batch to laya-ts (structured decisions over batched states).
- Follow both Agent convention (states, questions, opts) and Router convention (requests mapped with { state, questions } to predictBatch).
- Support minConfidence / min_confidence abstention gating on batched states, projecting low-confidence answers to null.
- Expose answer_confidence (calibrated max(p)) on DecisionResult, matching Python parity from PR NandhaKishorM#685.
- Expose decideBatch and decide_batch on Agent and Router.
- Add test coverage across schema projection, explicit questions, details, confidence gating, and error handling.
Angboo pushed a commit to Angboo/laya that referenced this pull request Sep 30, 2026
- Add decideBatch to laya-ts (structured decisions over batched states).
- Follow both Agent convention (states, questions, opts) and Router convention (requests mapped with { state, questions } to predictBatch).
- Support minConfidence abstention gating on batched states, projecting low-confidence answers to null.
- Expose answer_confidence (calibrated max(p)) on DecisionResult, matching Python parity from PR NandhaKishorM#685.
- Expose decideBatch on Agent and Router.
- Add test coverage across schema projection, explicit questions, details, confidence gating, and error handling.
Angboo pushed a commit to Angboo/laya that referenced this pull request Sep 30, 2026
- Add decideBatch to laya-ts (structured decisions over batched states).
- Follow both Agent convention (states, questions, opts) and Router convention (requests mapped with { state, questions } to predictBatch).
- Support minConfidence abstention gating on batched states, projecting low-confidence answers to null.
- Expose answer_confidence (calibrated max(p)) on DecisionResult, matching Python PR NandhaKishorM#685 parity.
- Expose decideBatch on Agent and Router.
- Add test coverage across schema projection, explicit questions, details, confidence gating, and error handling.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants