Skip to content

popoto.extraction.candidates

popoto.extraction.candidates

Deterministic candidate enumeration for auditable extraction (M3, #562).

A candidate is a span of one conversation turn that the rest of the extraction pipeline decides on: the firewall scans it, the verdict stage votes on it, and the decision log records exactly one terminal state for it.

This module is the pipeline's only source of candidates, and it is deliberately the dullest stage:

  • Pure. No Redis, no LLM, no network, no clock. generate_candidates is a function of (turn_id, text) alone.
  • Deterministic. The same input always yields the same list, in the same order, with the same ids. Auditability depends on this -- a non-deterministic candidate set makes a decision log unreplayable.
  • Exhaustive. Nothing is filtered here. Short sentences, duplicate sentences and low-value entities are all emitted, because dropping a candidate is a decision that must be logged by the caller, not a silence produced here. There is deliberately no per-turn cap.

Two generator rules produce the v1 candidate set:

sentence One candidate per sentence span, using the same split regex as :class:~popoto.extraction.HeuristicExtractionProvider. entity One candidate per pattern-lifted named entity. The lift is a regex, not a model call -- an LLM here would make the candidate set non-reproducible.

Candidate dataclass

One deterministically enumerated span of a turn.

Attributes:

Name Type Description
text str

The verbatim span. Byte-identical to turn_text[start:end] -- nothing is normalized, distilled or rewritten (that is M4's job).

turn_id str

Id of the turn this span came from.

candidate_id str

f"{turn_id}:{generator_rule}:{ordinal}". Unique within a turn, deterministic, and deliberately low-entropy: it is written to the provenance journal as a cand: subject tag, and the journal's write-time firewall blocks high-entropy tags such as hex digests. Never make this a hash.

start int

Character offset of the span's first character in the turn text.

end int

Character offset one past the span's last character.

generator_rule str

Which rule produced this candidate -- "sentence" or "entity".

Source code in src/popoto/extraction/candidates.py
@dataclass(frozen=True)
class Candidate:
    """One deterministically enumerated span of a turn.

    Attributes:
        text: The verbatim span. Byte-identical to
            ``turn_text[start:end]`` -- nothing is normalized, distilled or
            rewritten (that is M4's job).
        turn_id: Id of the turn this span came from.
        candidate_id: ``f"{turn_id}:{generator_rule}:{ordinal}"``. Unique
            within a turn, deterministic, and deliberately **low-entropy**:
            it is written to the provenance journal as a ``cand:`` subject
            tag, and the journal's write-time firewall blocks high-entropy
            tags such as hex digests. Never make this a hash.
        start: Character offset of the span's first character in the turn
            text.
        end: Character offset one past the span's last character.
        generator_rule: Which rule produced this candidate --
            ``"sentence"`` or ``"entity"``.
    """

    text: str
    turn_id: str
    candidate_id: str
    start: int
    end: int
    generator_rule: str

generate_candidates(turn_id, text)

Enumerate the complete v1 candidate set for one turn.

Parameters:

Name Type Description Default
turn_id str

Id of the turn being extracted. Becomes the first segment of every candidate_id.

required
text Optional[str]

The turn's raw text.

required

Returns:

Type Description
List[Candidate]

Sentence candidates in document order, followed by entity

List[Candidate]

candidates in document order. Empty list when text is empty,

List[Candidate]

whitespace-only or not a string -- an empty turn produces zero

List[Candidate]

candidates, and the caller (not this module) logs the

List[Candidate]

reject(empty_turn) row. This module has no decision-log

List[Candidate]

dependency.

Source code in src/popoto/extraction/candidates.py
def generate_candidates(turn_id: str, text: Optional[str]) -> List[Candidate]:
    """Enumerate the complete v1 candidate set for one turn.

    Args:
        turn_id: Id of the turn being extracted. Becomes the first segment
            of every ``candidate_id``.
        text: The turn's raw text.

    Returns:
        Sentence candidates in document order, followed by entity
        candidates in document order. Empty list when ``text`` is empty,
        whitespace-only or not a string -- an empty turn produces *zero*
        candidates, and the caller (not this module) logs the
        ``reject(empty_turn)`` row. This module has no decision-log
        dependency.
    """
    if not isinstance(text, str) or not text.strip():
        return []

    candidates = [
        _make(turn_id, text, span, SENTENCE_RULE, ordinal)
        for ordinal, span in enumerate(_sentence_spans(text))
    ]
    candidates += [
        _make(turn_id, text, span, ENTITY_RULE, ordinal)
        for ordinal, span in enumerate(_entity_spans(text))
    ]
    return candidates