Skip to content

popoto.extraction.verdict

popoto.extraction.verdict

Enum-confined per-candidate verdict stage for auditable extraction (M3).

This module owns the vocabulary the rest of the auditable-extraction pipeline decides in, and the one LLM call that produces a verdict for a single :class:~popoto.extraction.candidates.Candidate.

Two invariants make this stage auditable, and both are enforced here rather than trusted to the model:

  1. The model contributes enums only. A reply is parsed into exactly {candidate_id, verdict, reason_code}; verdict and reason_code must be members of the fixed vocabularies below, and candidate_id is taken from the trusted candidate, never from the reply. Any free text the model emits is discarded at the parse boundary, so no model-authored prose can reach the store. Accepted content is the verbatim candidate span, assembled by trusted code (see decision_log.py).
  2. No candidate is ever silently dropped. Every path through :func:llm_verdict returns a :class:VerdictResult -- a blocked candidate, a malformed reply, an unreachable API and a raising client all map to a logged verdict. The function never raises and never returns None.

The per-candidate never-record firewall (M2) runs before the LLM call, so text the firewall blocks is never transmitted to the provider.

Opt-in only: importing this module never imports anthropic. The optional dependency is resolved the same way popoto.extraction.claude resolves it, so callers and tests can always monkeypatch anthropic_module regardless of whether the package is installed.

Example

from popoto.extraction.candidates import generate_candidates from popoto.extraction.verdict import llm_verdict

for candidate in generate_candidates("turn-001", response_text): result = llm_verdict(candidate) # result.verdict / result.reason_code are enum members, always.

TERMINAL_VERDICTS = frozenset({Verdict.FIREWALL_DROP, Verdict.ACCEPT, Verdict.REJECT, Verdict.WITHHOLD}) module-attribute

The four terminal states. Verdict.PENDING is deliberately absent.

LLM_VERDICTS = frozenset({Verdict.ACCEPT, Verdict.REJECT, Verdict.WITHHOLD}) module-attribute

The only verdicts a model reply may carry.

FIREWALL_DROP is a trusted-code decision (the never-record firewall, before or after the call) and PENDING is an internal write-ordering marker; a reply claiming either is malformed.

LLM_REASON_CODES = {Verdict.ACCEPT: frozenset({ReasonCode.ACCEPTED}), Verdict.REJECT: frozenset({ReasonCode.NOT_A_FACT, ReasonCode.NOT_MEMORABLE}), Verdict.WITHHOLD: frozenset({ReasonCode.LOW_CONFIDENCE, ReasonCode.NEEDS_CONFIRMATION})} module-attribute

Reason codes the model may pair with each verdict it may emit.

A reply whose reason_code is absent from its verdict's set is malformed -- including a reason code that is a valid enum member but reserved for trusted code (e.g. assembly_failed). This is what keeps the model from labelling its own rejection as an infrastructure failure, which would corrupt the offline precision/recall breakdown.

VERDICT_MODEL = 'claude-haiku-4-5-20251001' module-attribute

Pinned model for the per-candidate verdict call. Not user-configurable.

Deliberately a smaller model than claude.py's EXTRACTION_MODEL: the verdict stage issues roughly one call per sentence-plus-entity candidate rather than one per turn (plan Risk 1, "LLM verdict-call cost per turn"), and the task is a constrained enum classification rather than open-ended extraction.

VERDICT_MAX_TOKENS = 256 module-attribute

Pinned max_tokens. The reply is three short enum-valued keys.

VERDICT_PROMPT = 'You are a memory-verdict engine for an AI agent. You are given ONE candidate span of text lifted verbatim from a conversation turn, and you decide whether it is worth remembering.\n\nReply with the candidate\'s id and exactly one verdict and one reason code:\n\n- "accept" (reason "accepted") -- a discrete, independently-useful fact worth remembering later.\n- "reject" with reason "not_a_fact" (it asserts nothing -- a greeting, a question, filler, conversational scaffolding) or "not_memorable" (it asserts something, but it is trivial or has no value once the conversation ends).\n- "withhold" with reason "low_confidence" (the span is hedged, speculative, or too ambiguous to store as stated) or "needs_confirmation" (it looks consequential but depends on context outside this span).\n\nYou are not asked to rewrite, summarize, or explain. Emit only the id, the verdict, and the reason code -- any other text is discarded.' module-attribute

Pinned system prompt for the verdict call. Not user-configurable.

VERDICT_SCHEMA = {'type': 'object', 'properties': {'candidate_id': {'type': 'string'}, 'verdict': {'type': 'string', 'enum': sorted(v.value for v in LLM_VERDICTS)}, 'reason_code': {'type': 'string', 'enum': sorted(r.value for codes in LLM_REASON_CODES.values() for r in codes)}}, 'required': ['candidate_id', 'verdict', 'reason_code'], 'additionalProperties': False} module-attribute

JSON schema confining the reply to enums. Not user-configurable.

The schema is a first line of defence, not the enforcement point: :func:_parse_reply re-validates every field, because a provider that ignores or partially honours the schema must still not be able to write free text or an out-of-vocabulary code into the decision log.

Verdict

Bases: str, Enum

The state vocabulary a candidate can be logged in.

Four of these are terminal: every candidate ends in exactly one of FIREWALL_DROP, ACCEPT, REJECT or WITHHOLD (issue

562's acceptance criteria).

PENDING is not terminal and is not a fifth terminal state. It is an intent marker written on the decision row before assembly calls the provenance journal, so a candidate can never reach an irreversible side effect with zero decision-log rows. A surviving PENDING row means a process died mid-assembly -- a visible, queryable, recoverable incident. Any terminal-state aggregation (per-turn summaries, offline precision/recall) must exclude it; use :data:TERMINAL_VERDICTS or :attr:Verdict.is_terminal rather than iterating the enum.

The model's own vocabulary is narrower still -- see :data:LLM_VERDICTS. The LLM never emits PENDING or FIREWALL_DROP; both are written exclusively by trusted code.

Source code in src/popoto/extraction/verdict.py
class Verdict(str, enum.Enum):
    """The state vocabulary a candidate can be logged in.

    Four of these are **terminal**: every candidate ends in exactly one of
    ``FIREWALL_DROP``, ``ACCEPT``, ``REJECT`` or ``WITHHOLD`` (issue
    #562's acceptance criteria).

    ``PENDING`` is **not terminal** and is not a fifth terminal state. It
    is an intent marker written on the decision row *before* assembly
    calls the provenance journal, so a candidate can never reach an
    irreversible side effect with zero decision-log rows. A surviving
    ``PENDING`` row means a process died mid-assembly -- a visible,
    queryable, recoverable incident. Any terminal-state aggregation
    (per-turn summaries, offline precision/recall) must exclude it; use
    :data:`TERMINAL_VERDICTS` or :attr:`Verdict.is_terminal` rather than
    iterating the enum.

    The model's own vocabulary is narrower still -- see
    :data:`LLM_VERDICTS`. The LLM never emits ``PENDING`` or
    ``FIREWALL_DROP``; both are written exclusively by trusted code.
    """

    FIREWALL_DROP = "firewall_drop"
    ACCEPT = "accept"
    REJECT = "reject"
    WITHHOLD = "withhold"
    PENDING = "pending"

    @property
    def is_terminal(self) -> bool:
        """True for the four terminal states, False for ``PENDING``."""
        return self is not Verdict.PENDING

is_terminal property

True for the four terminal states, False for PENDING.

ReasonCode

Bases: str, Enum

The fixed reason vocabulary paired with a :class:Verdict.

Only the members listed in :data:LLM_REASON_CODES may come from a model reply. The rest are written exclusively by trusted code:

  • PRE_LLM_CANDIDATE_BLOCK / POST_ACCEPT_JOURNAL_BLOCK -- never-record firewall refusals, before and after the LLM call respectively. Both pair with FIREWALL_DROP; the reason code is what distinguishes them, so FIREWALL_DROP keeps meaning exactly "privacy refusal" and never becomes a bucket for write errors.
  • ASSEMBLY_FAILED / AMBIGUOUS_RECONCILIATION -- assembly-time failures, both pairing with REJECT.
  • LLM_UNAVAILABLE -- the verdict call failed, returned nothing, or returned something that is not a well-formed enum reply. Pairs with REJECT. Offline analysis reads this as an infrastructure loss and must not charge it against the model's recall.
  • EMPTY_TURN -- there was nothing to decide (blank turn, or a candidate whose span is whitespace). Pairs with REJECT.
  • TURN_LEVEL_BLOCK -- the turn-level (M2) never-record scan voided the whole turn before any candidate was generated. Pairs with FIREWALL_DROP, but is distinct from PRE_LLM_CANDIDATE_BLOCK: that code means a per-candidate span was blocked by the M3 scan after candidates existed and the LLM never saw that one; this code means M2's turn-level scan fired first and no candidates were ever generated for this turn at all.
Source code in src/popoto/extraction/verdict.py
class ReasonCode(str, enum.Enum):
    """The fixed reason vocabulary paired with a :class:`Verdict`.

    Only the members listed in :data:`LLM_REASON_CODES` may come from a
    model reply. The rest are written exclusively by trusted code:

    - ``PRE_LLM_CANDIDATE_BLOCK`` / ``POST_ACCEPT_JOURNAL_BLOCK`` --
      never-record firewall refusals, before and after the LLM call
      respectively. Both pair with ``FIREWALL_DROP``; the reason code is
      what distinguishes them, so ``FIREWALL_DROP`` keeps meaning exactly
      "privacy refusal" and never becomes a bucket for write errors.
    - ``ASSEMBLY_FAILED`` / ``AMBIGUOUS_RECONCILIATION`` -- assembly-time
      failures, both pairing with ``REJECT``.
    - ``LLM_UNAVAILABLE`` -- the verdict call failed, returned nothing, or
      returned something that is not a well-formed enum reply. Pairs with
      ``REJECT``. Offline analysis reads this as an *infrastructure* loss
      and must not charge it against the model's recall.
    - ``EMPTY_TURN`` -- there was nothing to decide (blank turn, or a
      candidate whose span is whitespace). Pairs with ``REJECT``.
    - ``TURN_LEVEL_BLOCK`` -- the turn-level (M2) never-record scan voided
      the *whole* turn before any candidate was generated. Pairs with
      ``FIREWALL_DROP``, but is distinct from
      ``PRE_LLM_CANDIDATE_BLOCK``: that code means a per-candidate span was
      blocked by the M3 scan after candidates existed and the LLM never saw
      *that* one; this code means M2's turn-level scan fired first and no
      candidates were ever generated for this turn at all.
    """

    # --- trusted code only -------------------------------------------
    PRE_LLM_CANDIDATE_BLOCK = "pre_llm_candidate_block"
    POST_ACCEPT_JOURNAL_BLOCK = "post_accept_journal_block"
    ASSEMBLY_FAILED = "assembly_failed"
    AMBIGUOUS_RECONCILIATION = "ambiguous_reconciliation"
    LLM_UNAVAILABLE = "llm_unavailable"
    EMPTY_TURN = "empty_turn"
    TURN_LEVEL_BLOCK = "turn_level_block"

    # --- the model may emit these ------------------------------------
    ACCEPTED = "accepted"
    NOT_A_FACT = "not_a_fact"
    NOT_MEMORABLE = "not_memorable"
    LOW_CONFIDENCE = "low_confidence"
    NEEDS_CONFIRMATION = "needs_confirmation"

VerdictResult dataclass

One candidate's verdict: enums plus the trusted candidate id.

Deliberately has no free-text field. candidate_id is copied from the :class:~popoto.extraction.candidates.Candidate that trusted code generated -- it is never read out of the model's reply -- so nothing on this object originated as model prose.

Attributes:

Name Type Description
candidate_id str

The deciding candidate's id, from the candidate.

verdict Verdict

A :class:Verdict member.

reason_code ReasonCode

A :class:ReasonCode member.

Source code in src/popoto/extraction/verdict.py
@dataclass(frozen=True)
class VerdictResult:
    """One candidate's verdict: enums plus the trusted candidate id.

    Deliberately has no free-text field. ``candidate_id`` is copied from
    the :class:`~popoto.extraction.candidates.Candidate` that trusted code
    generated -- it is never read out of the model's reply -- so nothing
    on this object originated as model prose.

    Attributes:
        candidate_id: The deciding candidate's id, from the candidate.
        verdict: A :class:`Verdict` member.
        reason_code: A :class:`ReasonCode` member.
    """

    candidate_id: str
    verdict: Verdict
    reason_code: ReasonCode

llm_verdict(candidate, client=None)

Decide one candidate: firewall first, then a single LLM call.

Never raises and never returns None -- every failure mode maps to a logged verdict, because a candidate that vanishes is precisely the defect this module exists to eliminate.

Order of operations:

  1. A whitespace-only span is reject / empty_turn; there is nothing to decide and no call is made.
  2. scan_never_record(candidate.text) runs before the call. A blocked span is firewall_drop / pre_llm_candidate_block and its text is never transmitted to the provider.
  3. Otherwise one call is issued. A malformed, empty or out-of-vocabulary reply, an unreachable provider, and a raising client all map to reject / llm_unavailable.

Parameters:

Name Type Description Default
candidate Candidate

The candidate to decide. Only candidate_id and text are read.

required
client Any

An Anthropic-style client (anything exposing messages.create). None builds the default client, which requires the optional anthropic package.

None

Returns:

Name Type Description
A VerdictResult

class:VerdictResult carrying enum fields and the candidate's

VerdictResult

own id -- never any model-authored text.

Source code in src/popoto/extraction/verdict.py
def llm_verdict(candidate: "Candidate", client: Any = None) -> VerdictResult:
    """Decide one candidate: firewall first, then a single LLM call.

    Never raises and never returns ``None`` -- every failure mode maps to
    a logged verdict, because a candidate that vanishes is precisely the
    defect this module exists to eliminate.

    Order of operations:

    1. A whitespace-only span is ``reject`` / ``empty_turn``; there is
       nothing to decide and no call is made.
    2. ``scan_never_record(candidate.text)`` runs **before** the call. A
       blocked span is ``firewall_drop`` / ``pre_llm_candidate_block`` and
       its text is never transmitted to the provider.
    3. Otherwise one call is issued. A malformed, empty or
       out-of-vocabulary reply, an unreachable provider, and a raising
       client all map to ``reject`` / ``llm_unavailable``.

    Args:
        candidate: The candidate to decide. Only ``candidate_id`` and
            ``text`` are read.
        client: An Anthropic-style client (anything exposing
            ``messages.create``). ``None`` builds the default client,
            which requires the optional ``anthropic`` package.

    Returns:
        A :class:`VerdictResult` carrying enum fields and the candidate's
        own id -- never any model-authored text.
    """
    candidate_id = candidate.candidate_id

    if not candidate.text or not candidate.text.strip():
        logger.debug("verdict: blank candidate %s -> reject(empty_turn)", candidate_id)
        return VerdictResult(candidate_id, Verdict.REJECT, ReasonCode.EMPTY_TURN)

    firewall = scan_never_record(candidate.text)
    if firewall.blocked:
        # Log the reason code and the candidate id only -- never a
        # fragment of the blocked span.
        logger.info(
            "verdict: candidate %s blocked by never-record firewall (%s)",
            candidate_id,
            firewall.reason,
        )
        return VerdictResult(
            candidate_id,
            Verdict.FIREWALL_DROP,
            ReasonCode.PRE_LLM_CANDIDATE_BLOCK,
        )

    try:
        if client is None:
            client = _default_client()
        raw_text = _request_verdict(client, candidate)
    except Exception as e:
        logger.warning("verdict: call failed for candidate %s: %s", candidate_id, e)
        return VerdictResult(candidate_id, Verdict.REJECT, ReasonCode.LLM_UNAVAILABLE)

    result = _parse_reply(raw_text, candidate_id)
    if result is None:
        logger.warning(
            "verdict: malformed reply for candidate %s -> reject(llm_unavailable)",
            candidate_id,
        )
        return VerdictResult(candidate_id, Verdict.REJECT, ReasonCode.LLM_UNAVAILABLE)
    return result