Skip to content

popoto.extraction.resolution

popoto.extraction.resolution

Reference resolution stage for auditable extraction (M4, #563).

This module turns a candidate span plus its conversational window into a :class:Resolution: a rewritten statement with pronouns, relative dates and definite references either resolved, explicitly assumed, flagged as an evidence gap, or left indeterminate. Exactly four statuses make up the model's vocabulary, in :class:ResolutionStatus:

  • resolved -- the reference was anchored with no ambiguity.
  • assumed -- the stage anchored it, but only by stating an assumption (e.g. picking the most recent antecedent).
  • evidence_gap -- multiple plausible antecedents exist and the stage cannot pick one; the candidates are recorded and a clarifying question is posed.
  • indeterminate -- nothing in the window resolves the reference.

Fail-open, per M3's precedent. :func:resolve_references never raises and never returns None. A missing anthropic client, a raising client, a malformed reply, or the M4_RESOLUTION_ENABLED kill switch all fall back to a degraded :class:Resolution whose statement is byte-identical to the candidate's own verbatim text -- the entry is still captured. This is "quality loss, not corruption": M4 never causes M3's fail-open contract to fail closed.

The model never emits an epoch number. Every date the model returns is ISO-8601 text; this module is the only thing that ever calls datetime.fromisoformat on it, and floats reaching :class:Resolution are always Python-computed (see :func:_to_epoch), never model-supplied.

RESOLUTION_MODEL = 'claude-haiku-4-5-20251001' module-attribute

Pinned model for the per-candidate resolution call. Not user-configurable.

Same tier as verdict.py's VERDICT_MODEL -- this call runs once per accepted candidate, not once per turn.

RESOLUTION_MAX_TOKENS = 1024 module-attribute

Pinned max_tokens. A rewritten statement plus a handful of references.

RESOLUTION_PROMPT = 'You are a reference-resolution engine for an AI agent\'s memory. You are given ONE candidate statement lifted verbatim from a conversation turn, plus recent conversational context, and you resolve its pronouns, relative dates and definite references.\n\nFor each reference you find in the candidate text, report:\n- "surface": the EXACT substring of the candidate text it covers (character for character -- do not paraphrase it).\n- "start"/"end": its character offsets into the candidate text.\n- "kind": "pronoun", "relative_time", or "definite_reference".\n- "status": "resolved" (unambiguous), "assumed" (you picked one candidate and must state the assumption), "evidence_gap" (multiple plausible antecedents, list 2-4 of them and ask a clarifying question), or "indeterminate" (nothing in the context resolves it).\n- For a resolved or assumed reference: "resolved_text" (the resolved form) and, for "relative_time", "resolved_iso" (an ISO-8601 date/datetime string -- NEVER an epoch number).\n- For "relative_time" references only, also report "temporal_role": "onset" (the claim becomes true at this date), "deadline" (something is due by this date, but the claim is already true now), "mention" (the date is mentioned but is neither), or "none".\n- For "assumed": a one-line "assumption" explaining your choice.\n- For "evidence_gap": "candidates" (2-4 plausible antecedents) and a "question" that would resolve the ambiguity.\n\nThen report "statement": the candidate rewritten with resolved references substituted in place, changing nothing else. If nothing needed resolving, "references" is an empty array and "statement" equals the candidate text verbatim.\n\nReply with the candidate\'s id, the statement, and the references array only -- any other text is discarded.' module-attribute

Pinned system prompt for the resolution call. Not user-configurable.

RESOLUTION_SCHEMA = {'type': 'object', 'properties': {'candidate_id': {'type': 'string'}, 'statement': {'type': 'string'}, 'references': {'type': 'array', 'items': {'type': 'object', 'properties': {'surface': {'type': 'string'}, 'start': {'type': 'integer'}, 'end': {'type': 'integer'}, 'kind': {'type': 'string', 'enum': [k.value for k in ReferenceKind]}, 'status': {'type': 'string', 'enum': [s.value for s in ResolutionStatus]}, 'temporal_role': {'type': 'string', 'enum': [r.value for r in TemporalRole]}, 'resolved_text': {'type': ['string', 'null']}, 'resolved_iso': {'type': ['string', 'null']}, 'assumption': {'type': ['string', 'null']}, 'candidates': {'type': 'array', 'items': {'type': 'string'}}, 'question': {'type': ['string', 'null']}}, 'required': ['surface', 'start', 'end', 'kind', 'status', 'temporal_role', 'resolved_text', 'resolved_iso', 'assumption', 'candidates', 'question'], 'additionalProperties': False}}}, 'required': ['candidate_id', 'statement', 'references'], 'additionalProperties': False} module-attribute

JSON schema confining the reply. Not user-configurable.

One level of nesting only -- an object holding a flat array of flat objects -- because Anthropic's structured-output schema support is not full JSON Schema. This is a first line of defence, not the enforcement point: :func:_parse_reply re-validates every field.

ResolutionStatus

Bases: str, Enum

The four-member status vocabulary a resolved reference may carry.

Ordered worst-last for :func:ResolutionStatus.worst_of: resolved is the best outcome, indeterminate the worst.

Source code in src/popoto/extraction/resolution.py
class ResolutionStatus(str, enum.Enum):
    """The four-member status vocabulary a resolved reference may carry.

    Ordered worst-last for :func:`ResolutionStatus.worst_of`:
    ``resolved`` is the best outcome, ``indeterminate`` the worst.
    """

    RESOLVED = "resolved"
    ASSUMED = "assumed"
    EVIDENCE_GAP = "evidence_gap"
    INDETERMINATE = "indeterminate"

    @property
    def _rank(self) -> int:
        return _STATUS_RANK[self]

    @staticmethod
    def worst_of(statuses: "List[ResolutionStatus]") -> "ResolutionStatus":
        """Return the worst (most-degraded) status among ``statuses``.

        Defaults to ``RESOLVED`` (the best status) for an empty input --
        this mirrors "nothing needed resolving" being a legitimate,
        non-degraded outcome (see :func:`resolve_references`).
        """
        if not statuses:
            return ResolutionStatus.RESOLVED
        return max(statuses, key=lambda s: s._rank)

worst_of(statuses) staticmethod

Return the worst (most-degraded) status among statuses.

Defaults to RESOLVED (the best status) for an empty input -- this mirrors "nothing needed resolving" being a legitimate, non-degraded outcome (see :func:resolve_references).

Source code in src/popoto/extraction/resolution.py
@staticmethod
def worst_of(statuses: "List[ResolutionStatus]") -> "ResolutionStatus":
    """Return the worst (most-degraded) status among ``statuses``.

    Defaults to ``RESOLVED`` (the best status) for an empty input --
    this mirrors "nothing needed resolving" being a legitimate,
    non-degraded outcome (see :func:`resolve_references`).
    """
    if not statuses:
        return ResolutionStatus.RESOLVED
    return max(statuses, key=lambda s: s._rank)

ReferenceKind

Bases: str, Enum

What kind of reference a span is.

Source code in src/popoto/extraction/resolution.py
class ReferenceKind(str, enum.Enum):
    """What kind of reference a span is."""

    PRONOUN = "pronoun"
    RELATIVE_TIME = "relative_time"
    DEFINITE_REFERENCE = "definite_reference"

TemporalRole

Bases: str, Enum

The role a relative_time reference plays in the statement.

Only onset participates in the valid_from emission rule (see :func:_compute_valid_from); deadline and mention reference a date without asserting the claim became true then, and none is for non-temporal references classified with this vocabulary by mistake-proofing rather than by relevance.

Source code in src/popoto/extraction/resolution.py
class TemporalRole(str, enum.Enum):
    """The role a ``relative_time`` reference plays in the statement.

    Only ``onset`` participates in the ``valid_from`` emission rule (see
    :func:`_compute_valid_from`); ``deadline`` and ``mention`` reference
    a date without asserting the claim became true then, and ``none`` is
    for non-temporal references classified with this vocabulary by
    mistake-proofing rather than by relevance.
    """

    ONSET = "onset"
    DEADLINE = "deadline"
    MENTION = "mention"
    NONE = "none"

Reference dataclass

One resolved (or attempted) reference inside a candidate span.

Attributes:

Name Type Description
surface str

The exact substring of the candidate text this reference covers -- candidate.text[start:end] == surface is mechanically enforced by :func:_parse_reference.

start int

Start offset into the candidate text.

end int

End offset into the candidate text.

kind ReferenceKind

What kind of reference this is.

status ResolutionStatus

This reference's own resolution status.

temporal_role TemporalRole

Only meaningful when kind == RELATIVE_TIME.

resolved_text Optional[str]

The resolved human-readable text, when resolved or assumed.

resolved_epoch Optional[float]

The resolved instant as a Python-computed epoch float (via :func:_to_epoch), or None. Never taken directly from the model -- see the module docstring.

assumption Optional[str]

A one-line stated assumption, required when status == ASSUMED.

candidates Tuple[str, ...]

Plausible antecedents, required (2-4) when status == EVIDENCE_GAP.

question Optional[str]

A clarifying question, required when status == EVIDENCE_GAP.

Source code in src/popoto/extraction/resolution.py
@dataclass(frozen=True)
class Reference:
    """One resolved (or attempted) reference inside a candidate span.

    Attributes:
        surface: The exact substring of the candidate text this reference
            covers -- ``candidate.text[start:end] == surface`` is
            mechanically enforced by :func:`_parse_reference`.
        start: Start offset into the candidate text.
        end: End offset into the candidate text.
        kind: What kind of reference this is.
        status: This reference's own resolution status.
        temporal_role: Only meaningful when ``kind == RELATIVE_TIME``.
        resolved_text: The resolved human-readable text, when resolved or
            assumed.
        resolved_epoch: The resolved instant as a Python-computed epoch
            float (via :func:`_to_epoch`), or ``None``. Never taken
            directly from the model -- see the module docstring.
        assumption: A one-line stated assumption, required when
            ``status == ASSUMED``.
        candidates: Plausible antecedents, required (2-4) when
            ``status == EVIDENCE_GAP``.
        question: A clarifying question, required when
            ``status == EVIDENCE_GAP``.
    """

    surface: str
    start: int
    end: int
    kind: ReferenceKind
    status: ResolutionStatus
    temporal_role: TemporalRole = TemporalRole.NONE
    resolved_text: Optional[str] = None
    resolved_epoch: Optional[float] = None
    assumption: Optional[str] = None
    candidates: Tuple[str, ...] = ()
    question: Optional[str] = None

WindowTurn dataclass

One prior turn carried in a :class:TurnContext window.

Attributes:

Name Type Description
turn_id Optional[str]

Identity of the turn.

speaker Optional[str]

Who spoke it, or None if unknown.

text str

The turn's text.

Source code in src/popoto/extraction/resolution.py
@dataclass(frozen=True)
class WindowTurn:
    """One prior turn carried in a :class:`TurnContext` window.

    Attributes:
        turn_id: Identity of the turn.
        speaker: Who spoke it, or ``None`` if unknown.
        text: The turn's text.
    """

    turn_id: Optional[str]
    speaker: Optional[str]
    text: str

TurnContext dataclass

The conversational context a candidate is resolved against.

Attributes:

Name Type Description
speaker Optional[str]

Who spoke the current turn, or None if unknown.

captured_at float

Epoch seconds the current turn was captured. A None, NaN or infinite value is coerced to time.time() in __post_init__ (with a logged warning) -- this is a guard M4 adds, not M1 behaviour; see the module docstring's fail-open discussion and Technical Approach §3a of the plan.

timezone str

IANA timezone name naive resolved dates are anchored to.

window Tuple[WindowTurn, ...]

Prior turns, oldest-first, that make up the resolution window. Use :meth:bounded_window rather than reading this directly when a bound is required.

Source code in src/popoto/extraction/resolution.py
@dataclass(frozen=True)
class TurnContext:
    """The conversational context a candidate is resolved against.

    Attributes:
        speaker: Who spoke the current turn, or ``None`` if unknown.
        captured_at: Epoch seconds the current turn was captured. A
            ``None``, NaN or infinite value is coerced to ``time.time()``
            in ``__post_init__`` (with a logged warning) -- this is a
            guard M4 adds, not M1 behaviour; see the module docstring's
            fail-open discussion and Technical Approach §3a of the plan.
        timezone: IANA timezone name naive resolved dates are anchored
            to.
        window: Prior turns, oldest-first, that make up the resolution
            window. Use :meth:`bounded_window` rather than reading this
            directly when a bound is required.
    """

    speaker: Optional[str] = None
    captured_at: float = field(default_factory=time.time)
    timezone: str = "UTC"
    window: Tuple[WindowTurn, ...] = ()

    def __post_init__(self) -> None:
        captured_at = self.captured_at
        if captured_at is None or not math.isfinite(captured_at):
            logger.warning(
                "resolution: TurnContext.captured_at %r is missing or "
                "non-finite; coercing to the current clock",
                captured_at,
            )
            # frozen dataclass -- object.__setattr__ is the sanctioned
            # escape hatch for __post_init__ coercion.
            object.__setattr__(self, "captured_at", time.time())

    @classmethod
    def now(cls) -> "TurnContext":
        """Build a context with no speaker, no window, UTC, now."""
        return cls(speaker=None, captured_at=time.time(), timezone="UTC", window=())

    def bounded_window(self) -> Tuple[Tuple[WindowTurn, ...], bool]:
        """Return the window truncated to the M4 bounds, oldest-first.

        Truncation drops the *oldest* turns first when either the turn
        count or the total character count exceeds its bound (Technical
        Approach §9). Returns ``(turns, truncated)`` so callers can
        record whether truncation happened, distinguishing "the window
        did not contain the antecedent" from "the model missed it".
        """
        from ..fields.constants import Defaults

        max_turns = Defaults.M4_WINDOW_MAX_TURNS
        max_chars = Defaults.M4_WINDOW_MAX_CHARS

        turns = list(self.window)
        original_len = len(turns)

        if len(turns) > max_turns:
            turns = turns[-max_turns:]

        total_chars = sum(len(t.text) for t in turns)
        while turns and total_chars > max_chars:
            dropped = turns.pop(0)
            total_chars -= len(dropped.text)

        truncated = len(turns) < original_len
        return tuple(turns), truncated

now() classmethod

Build a context with no speaker, no window, UTC, now.

Source code in src/popoto/extraction/resolution.py
@classmethod
def now(cls) -> "TurnContext":
    """Build a context with no speaker, no window, UTC, now."""
    return cls(speaker=None, captured_at=time.time(), timezone="UTC", window=())

bounded_window()

Return the window truncated to the M4 bounds, oldest-first.

Truncation drops the oldest turns first when either the turn count or the total character count exceeds its bound (Technical Approach §9). Returns (turns, truncated) so callers can record whether truncation happened, distinguishing "the window did not contain the antecedent" from "the model missed it".

Source code in src/popoto/extraction/resolution.py
def bounded_window(self) -> Tuple[Tuple[WindowTurn, ...], bool]:
    """Return the window truncated to the M4 bounds, oldest-first.

    Truncation drops the *oldest* turns first when either the turn
    count or the total character count exceeds its bound (Technical
    Approach §9). Returns ``(turns, truncated)`` so callers can
    record whether truncation happened, distinguishing "the window
    did not contain the antecedent" from "the model missed it".
    """
    from ..fields.constants import Defaults

    max_turns = Defaults.M4_WINDOW_MAX_TURNS
    max_chars = Defaults.M4_WINDOW_MAX_CHARS

    turns = list(self.window)
    original_len = len(turns)

    if len(turns) > max_turns:
        turns = turns[-max_turns:]

    total_chars = sum(len(t.text) for t in turns)
    while turns and total_chars > max_chars:
        dropped = turns.pop(0)
        total_chars -= len(dropped.text)

    truncated = len(turns) < original_len
    return tuple(turns), truncated

Resolution dataclass

The outcome of resolving one candidate against its context.

Attributes:

Name Type Description
statement str

The rewritten statement. Byte-identical to verbatim when the stage is degraded, or when the model emitted an empty references array (nothing needed resolving).

verbatim str

The original candidate text, unmodified.

references Tuple[Reference, ...]

The resolved (or dropped-then-retried) references, in offset order.

status ResolutionStatus

The aggregate :class:ResolutionStatus -- the worst status among references, or RESOLVED when references is empty.

valid_from Optional[float]

The onset instant, as a Python-computed epoch float, or None. See the module-level onset rule documented on :func:_compute_valid_from.

degraded bool

True when this Resolution is a fail-open fallback (missing dependency, raising client, malformed reply, kill switch, or a non-finite float caught before construction). Distinct from status == INDETERMINATE, which can also be a genuine model abstention -- see :attr:subject_tag.

context Optional[TurnContext]

The :class:TurnContext this resolution ran against, or None.

window_truncated bool

Whether :meth:TurnContext.bounded_window reported truncation for this run.

Source code in src/popoto/extraction/resolution.py
@dataclass(frozen=True)
class Resolution:
    """The outcome of resolving one candidate against its context.

    Attributes:
        statement: The rewritten statement. Byte-identical to
            ``verbatim`` when the stage is degraded, or when the model
            emitted an empty references array (nothing needed
            resolving).
        verbatim: The original candidate text, unmodified.
        references: The resolved (or dropped-then-retried) references,
            in offset order.
        status: The aggregate :class:`ResolutionStatus` -- the worst
            status among ``references``, or ``RESOLVED`` when
            ``references`` is empty.
        valid_from: The onset instant, as a Python-computed epoch float,
            or ``None``. See the module-level onset rule documented on
            :func:`_compute_valid_from`.
        degraded: ``True`` when this Resolution is a fail-open fallback
            (missing dependency, raising client, malformed reply, kill
            switch, or a non-finite float caught before construction).
            Distinct from ``status == INDETERMINATE``, which can also be
            a genuine model abstention -- see :attr:`subject_tag`.
        context: The :class:`TurnContext` this resolution ran against,
            or ``None``.
        window_truncated: Whether :meth:`TurnContext.bounded_window`
            reported truncation for this run.
    """

    statement: str
    verbatim: str
    references: Tuple[Reference, ...] = ()
    status: ResolutionStatus = ResolutionStatus.RESOLVED
    valid_from: Optional[float] = None
    degraded: bool = False
    context: Optional[TurnContext] = None
    window_truncated: bool = False

    @property
    def subject_tag(self) -> str:
        """The ``res:*`` journal subject tag for this resolution.

        ``degraded`` takes precedence: a degraded resolution is tagged
        ``res:degraded`` and nothing else, so "the model abstained" and
        "the resolution stage never ran" remain distinguishable on the
        one channel guaranteed to travel with the fact (Technical
        Approach §3b).
        """
        if self.degraded:
            return "res:degraded"
        return f"res:{self.status.value}"

subject_tag property

The res:* journal subject tag for this resolution.

degraded takes precedence: a degraded resolution is tagged res:degraded and nothing else, so "the model abstained" and "the resolution stage never ran" remain distinguishable on the one channel guaranteed to travel with the fact (Technical Approach §3b).

resolve_references(candidate, turn_text, context, client=None)

Resolve one candidate's pronouns, dates and definite references.

Never raises and never returns None. Every failure mode -- Defaults.M4_RESOLUTION_ENABLED is False, the optional anthropic package is unavailable, the client raises, or the reply is malformed -- maps to the degraded fallback: Resolution(statement=verbatim, status=INDETERMINATE, references=(), valid_from=None, degraded=True), matching M3's fail-open contract (quality loss, not corruption).

An empty/whitespace candidate span is also degraded, checked here as defence in depth even though M3 rejects such candidates before they can reach this stage -- no client call is made in that case either.

A reply whose references array is empty is a legitimate, non-degraded outcome ("nothing needed resolving"): status = RESOLVED, statement == verbatim, no valid_from.

Parameters:

Name Type Description Default
candidate Candidate

The candidate to resolve. Only candidate_id and text are read.

required
turn_text str

The full turn text the candidate was lifted from.

required
context TurnContext

The conversational context to resolve against.

required
client Any

An Anthropic-style client (anything exposing messages.create). None builds the default client, which requires the optional anthropic package.

None

Returns:

Name Type Description
A Resolution

class:Resolution, always -- degraded or not.

Source code in src/popoto/extraction/resolution.py
def resolve_references(
    candidate: "Candidate",
    turn_text: str,
    context: TurnContext,
    client: Any = None,
) -> Resolution:
    """Resolve one candidate's pronouns, dates and definite references.

    Never raises and never returns ``None``. Every failure mode --
    ``Defaults.M4_RESOLUTION_ENABLED`` is False, the optional
    ``anthropic`` package is unavailable, the client raises, or the reply
    is malformed -- maps to the degraded fallback:
    ``Resolution(statement=verbatim, status=INDETERMINATE, references=(),
    valid_from=None, degraded=True)``, matching M3's fail-open contract
    (quality loss, not corruption).

    An empty/whitespace candidate span is also degraded, checked here as
    defence in depth even though M3 rejects such candidates before they
    can reach this stage -- no client call is made in that case either.

    A reply whose ``references`` array is empty is a *legitimate,
    non-degraded* outcome ("nothing needed resolving"): ``status =
    RESOLVED``, ``statement == verbatim``, no ``valid_from``.

    Args:
        candidate: The candidate to resolve. Only ``candidate_id`` and
            ``text`` are read.
        turn_text: The full turn text the candidate was lifted from.
        context: The conversational context to resolve against.
        client: An Anthropic-style client (anything exposing
            ``messages.create``). ``None`` builds the default client,
            which requires the optional ``anthropic`` package.

    Returns:
        A :class:`Resolution`, always -- degraded or not.
    """
    from ..fields.constants import Defaults

    candidate_id = candidate.candidate_id

    if not Defaults.M4_RESOLUTION_ENABLED:
        logger.debug(
            "resolution: M4_RESOLUTION_ENABLED is False; degrading %s",
            candidate_id,
        )
        return _degraded_resolution(candidate, context)

    if not candidate.text or not candidate.text.strip():
        logger.debug("resolution: blank candidate %s; degrading", candidate_id)
        return _degraded_resolution(candidate, context)

    window, truncated = context.bounded_window()
    bounded_context = TurnContext(
        speaker=context.speaker,
        captured_at=context.captured_at,
        timezone=context.timezone,
        window=window,
    )

    try:
        if client is None:
            client = _default_client()
        raw_text = _request_resolution(client, candidate, turn_text, bounded_context)
    except Exception as e:
        logger.warning("resolution: call failed for candidate %s: %s", candidate_id, e)
        return dataclasses.replace(
            _degraded_resolution(candidate, context), window_truncated=truncated
        )

    result = _parse_reply(raw_text, candidate, bounded_context.timezone)
    if result is None:
        logger.warning(
            "resolution: malformed reply for candidate %s; degrading",
            candidate_id,
        )
        return dataclasses.replace(
            _degraded_resolution(candidate, context), window_truncated=truncated
        )

    return dataclasses.replace(result, context=context, window_truncated=truncated)