popoto.extraction.decision_log¶
popoto.extraction.decision_log
¶
The write-side audit store for auditable extraction (M3).
Every candidate the generator produces gets a row here, and every row ends
in exactly one terminal state -- firewall_drop | accept | reject |
withhold (issue #562). Extraction precision and recall become computable
offline from these rows alone: "audit by construction".
Two properties carry the whole design, and both are structural rather than conventional -- neither depends on a caller remembering to do the right thing.
1. The row is identified by the candidate, not by a mint-on-save id.
agent_id, turn_id and candidate_id are all KeyFields, so
the composite Redis key is the candidate's identity and re-saving the
same tuple transitions that row in place. Do not "fix" this by copying
the sibling :class:~popoto.recipes.provenance_journal.JournalEntry, which
uses entry_id = AutoKeyField(): an AutoKeyField mints a brand-new
row on every save, which is right for an append-only journal and fatal
here. It would leave two rows behind every pending -> terminal
transition, make the terminal-write guard read an empty row and never
refuse, and make list_pending return already-reconciled rows forever.
AutoKeyField is forbidden on :class:DecisionRecord.
Two consequences of composite keying worth knowing before you read a key by eye or by position:
- KeyFields join alphabetically, not in declaration order
(
models/base.py:284-301), so the Redis key isDecisionRecord:<agent_id>:<candidate_id>:<turn_id>--candidate_idin the middle. DB_keyescapes colons and glob characters inside values (models/db_key.py:86-88), so acandidate_idoft-41:sent:0renders ast/-41{:}sent{:}0. That is correct; do not "fix" it.
2. The decision log is written before every irreversible side effect. A candidate never reaches a side effect that has no row already describing it:
firewall_drop/reject/withholdhave no downstream side effect, so their first write is also their terminal write.acceptis the only two-phase path: :meth:DecisionLog.write_pendingcommits a non-terminalpendingrow before assembly callsProvenanceJournal.append(), and :meth:DecisionLog.write_terminaltransitions that same row afterwards.pendingis not a fifth terminal state -- it is a visible unfinished write, which is the opposite of a silent drop. A survivingpendingrow means a process died mid-assembly.
Recovering stale pending rows is manual in v1 -- nothing sweeps,
expires or alerts on them, and decision rows deliberately carry no TTL
(a TTL would delete the audit evidence this module exists to keep). The
operator recipe is:
log = DecisionLog()
for row in log.list_pending("agent-7"): # oldest-first
print(row.turn_id, row.candidate_id, row.written_at)
# then re-invoke the auditable path for each stale (agent_id, turn_id);
# assembly's identity probe reconciles it.
A periodic age-keyed sweep, an alert threshold and any dashboard over this
log are M9 (#568) follow-ons, not v1. Detail rows ship unbounded on
purpose: the right retention horizon cannot be known until M9 consumes the
log, so v1 declines to guess a number. There is no LTRIM here and no
Defaults cap constant for this log.
CLAIM_KEY_PREFIX = 'popoto:m3:claim'
module-attribute
¶
Prefix for the ephemeral assembly-claim keys.
A separate, extraction-owned keyspace from the decision rows. Claim keys carry a TTL and hold no audit content, so losing one costs at most a reprobe -- never a record.
CLAIM_RELEASE_LUA = "\n-- KEYS[1] = the claim key, ARGV[1] = this runner's token\nif redis.call('GET', KEYS[1]) == ARGV[1] then\n return redis.call('DEL', KEYS[1])\nend\nreturn 0\n"
module-attribute
¶
Token-checked release. A runner never deletes a claim it lost.
SUMMARY_KEY_PREFIX = 'popoto:m3:summary'
module-attribute
¶
Prefix for the per-turn compact summary hash.
The summary is a convenience index, not a completeness fallback -- detail rows are unbounded and remain the sole source of truth. It exists so per-turn terminal-state counts are O(1) instead of a scan over every candidate row for the turn. Any summary/detail disagreement is resolved in favour of the detail rows.
DecisionRecord
¶
Bases: Model
One candidate's decision row, keyed by the candidate's identity.
agent_id + turn_id + candidate_id form the composite key,
so writing the same tuple twice transitions one row rather than
creating two. See the module docstring for why AutoKeyField is
forbidden here and for the alphabetical key-join order.
Every field below is a plain (non-indexed) field on purpose. Popoto
routes IndexedFieldMixin fields out of the base HSET and into
INDEX_SWAP_LUA (models/base.py:1470-1480), so an indexed
state would be invisible to the guard script's HGET and its
index would be silently desynchronised by the guard's HSET.
list_pending filters in Python instead; it is an operator recovery
reader, not a hot path.
Attributes:
| Name | Type | Description |
|---|---|---|
agent_id |
Owning agent. KeyField. |
|
turn_id |
The turn the candidate was generated from. KeyField. |
|
candidate_id |
|
|
state |
A :class: |
|
reason_code |
A :class: |
|
generator_rule |
Which deterministic rule produced the candidate, so offline metrics can break results down per rule. |
|
span_start |
Candidate span start offset in the turn text. |
|
span_end |
Candidate span end offset in the turn text. |
|
text_hash |
SHA-256 of the candidate text. A digest is safe here
because nothing scans a decision row; the |
|
entry_id |
The journal entry id, set on a terminal |
|
detail_code |
Free-form diagnostic string, written only by trusted
code. See :attr: |
|
written_at |
Unix timestamp, stamped on every write -- both the
|
Source code in src/popoto/extraction/decision_log.py
is_terminal
property
¶
True when :attr:state is one of the four terminal states.
DecisionLog
¶
Writer/reader over :class:DecisionRecord rows.
Stateless -- every method takes the identity it operates on -- so one instance can serve any agent. All writes go through Redis/Valkey core commands and Lua only.
Example::
log = DecisionLog()
log.write_terminal(
agent_id="agent-7",
candidate=candidate,
state=Verdict.FIREWALL_DROP,
reason_code=ReasonCode.PRE_LLM_CANDIDATE_BLOCK,
)
Source code in src/popoto/extraction/decision_log.py
300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 742 743 744 745 746 747 748 749 750 751 752 753 754 755 756 757 758 759 760 761 762 763 764 765 766 767 768 769 770 771 772 773 774 775 776 777 778 779 780 781 782 783 784 785 786 787 788 789 790 791 792 793 794 795 796 797 798 799 800 801 802 803 804 805 806 807 808 809 810 811 812 813 814 815 816 817 818 819 820 821 822 823 824 825 826 827 828 829 830 831 832 833 834 835 836 837 838 839 840 841 842 843 844 845 846 847 848 849 850 851 852 853 854 855 856 857 858 859 860 861 862 863 864 865 866 867 868 869 870 871 872 873 874 875 876 877 878 879 880 881 882 883 884 885 886 887 888 889 890 891 892 893 894 895 896 897 898 899 900 901 902 903 904 905 906 907 908 909 910 911 912 913 914 915 916 917 918 919 920 921 922 923 924 925 926 927 928 929 930 931 932 933 934 935 936 937 938 939 940 941 942 943 944 945 946 947 948 949 950 951 952 953 954 955 956 957 958 959 960 961 962 963 964 965 966 967 968 969 970 971 972 973 974 975 976 977 978 979 980 981 982 983 984 985 986 987 988 989 990 991 992 993 994 995 996 997 998 999 1000 1001 1002 1003 1004 1005 1006 1007 1008 1009 1010 1011 1012 1013 1014 1015 1016 1017 1018 1019 1020 1021 1022 1023 1024 1025 1026 1027 1028 1029 1030 1031 1032 1033 1034 1035 1036 1037 1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 | |
CONFLICT_REFUSED = 'terminal_conflict_refused'
class-attribute
instance-attribute
¶
detail_code recorded when a terminal write is refused.
detail_code is a free-form diagnostic string, not an enum: it
carries three structurally different payloads by design -- this fixed
literal, an exception class name on assembly_failed, and
",".join(entry_ids) on ambiguous_reconciliation. That does not
weaken the enums-only constraint on the model's output: state and
reason_code are genuine single-value enums, and detail_code is
written exclusively by trusted code, never by the LLM.
row_key(agent_id, turn_id, candidate_id)
staticmethod
¶
The Redis key for one candidate's row.
Built through DB_key rather than by string formatting, so the
alphabetical join order and the colon/glob escaping stay in one
place (the model) instead of being re-derived by every caller.
Source code in src/popoto/extraction/decision_log.py
summary_key(agent_id, turn_id)
staticmethod
¶
claim_key(agent_id, turn_id, candidate_id)
staticmethod
¶
The ephemeral assembly-claim key for one candidate.
Carries a TTL and holds no audit content -- losing it costs at most a re-probe, never a record.
Source code in src/popoto/extraction/decision_log.py
key_field_index_keys(agent_id, turn_id, candidate_id)
staticmethod
¶
The KeyField secondary index Sets for one row's key values.
Built with KeyFieldMixin's own key builder rather than by
formatting the $KeyF: pattern here, so the guard script SADDs
into exactly the Sets on_save would have.
Source code in src/popoto/extraction/decision_log.py
write_pending(agent_id, candidate, reason_code=ReasonCode.ACCEPTED)
¶
Commit the non-terminal pending row for an accepted candidate.
Phase 1 of the two-phase accept path. This write is committed
before ProvenanceJournal.append() is called, never
pipelined with it -- that ordering is what guarantees no candidate
can reach an irreversible side effect with zero decision-log rows.
No summary update accompanies it: the per-turn summary aggregates
terminal states only, and pending is never counted into it.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_id
|
str
|
Owning agent. Never |
required |
candidate
|
Candidate
|
The candidate being assembled. |
required |
reason_code
|
Union[ReasonCode, str]
|
The verdict's reason code, carried forward. |
ACCEPTED
|
Returns:
| Type | Description |
|---|---|
DecisionRecord
|
The saved :class: |
Source code in src/popoto/extraction/decision_log.py
write_terminal(agent_id, candidate, state, reason_code, entry_id='', detail_code='')
¶
Write a terminal state through the guarded script.
This is the only way any terminal state is written. Every
terminal state routes through it, including the pre-LLM
firewall_drop that cannot conflict in practice -- there is no
fast path to bypass.
The guard applies uniformly rather than only to non-accept
writes. Applying it to accept too costs nothing and removes the
last conditional branch: the legitimate pending -> accept
transition is never refused (a pending row is not accept
with an entry_id), and the only accept write it does refuse
is a duplicate assembly of an already-assembled candidate, where
refusing is correct.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_id
|
str
|
Owning agent. |
required |
candidate
|
Candidate
|
The candidate being decided. |
required |
state
|
Union[Verdict, str]
|
A terminal :class: |
required |
reason_code
|
Union[ReasonCode, str]
|
A :class: |
required |
entry_id
|
str
|
Journal entry id. Required on |
''
|
detail_code
|
str
|
Free-form trusted diagnostic string. |
''
|
Returns:
| Type | Description |
|---|---|
bool
|
|
bool
|
refused because the row is already terminal |
bool
|
|
bool
|
existing row stands and its |
bool
|
attr: |
Raises:
| Type | Description |
|---|---|
ValueError
|
If |
Source code in src/popoto/extraction/decision_log.py
441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 | |
acquire_claim(agent_id, turn_id, candidate_id)
¶
Claim a candidate for assembly. One atomic op, not read-then-act.
Without this, two runners over the same candidate -- a duplicated
delivery racing a crash-retry, both seeing a surviving pending
-- can both probe, both find nothing (neither has committed
append() yet), and both append. That produces two permanent
journal entries, and the journal cannot catch it: append()
takes no idempotency key and AppendOnlyViolation fires only when
a record's Redis key already exists, which is never true for a fresh
AutoKeyField append.
SET ... NX PX is one round trip and a core command on both Redis
and Valkey. It is chosen over WATCH/MULTI compare-and-set
deliberately: WATCH needs a dedicated connection held across the
transaction plus a retry loop, which does not compose with Popoto's
shared pool. Both are Valkey-safe; SET NX is smaller.
Returns:
| Type | Description |
|---|---|
Optional[str]
|
This runner's token if the claim was won, else |
Optional[str]
|
runner holding |
Optional[str]
|
no row transition -- it no-ops and leaves the candidate to |
Optional[str]
|
the winner. |
Source code in src/popoto/extraction/decision_log.py
release_claim(agent_id, turn_id, candidate_id, token)
¶
Release a claim this runner still owns.
Token-checked so a runner can never delete a claim it no longer owns -- if the TTL expired and another runner re-claimed, the DEL would otherwise hand that runner's candidate to a third.
Source code in src/popoto/extraction/decision_log.py
assemble(agent_id, candidate, journal, speaker=None, topic_tags=(), resolution=None)
¶
Assemble one accepted candidate into the provenance journal.
The full ordering, which is the reason this is one function rather than a few helpers a caller sequences by hand: claim -> read row -> (probe) -> pending -> append -> terminal transition -> release.
Dedup is owned here and cannot be delegated to the journal:
append() accepts no idempotency key, JournalEntry.entry_id
is an AutoKeyField, and M1 is append-only with no delete path,
so a duplicate entry would be permanent.
The four-case probe on the existing row:
- terminal
acceptwith anentry_id-- already assembled, skip entirely. - any other terminal state -- already decided, skip.
pending-- a retry of an interrupted assembly. Reconcile by candidate identity (see below) before considering a re-append.- absent -- fresh candidate. Write
pending, thenappend().
Reconciliation matches the cand:{candidate_id} subject tag, not
verbatim text. Text matching is unsound here and the plan
withdrew it: verbatim is not unique per candidate within a turn
by construction -- a repeated sentence, or a sentence span whose
text equals an entity-lifted span, produces two candidates with
identical verbatim, and a pending row could reconcile onto
the other candidate's entry and record the wrong entry_id.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_id
|
str
|
Owning agent. Required and never |
required |
candidate
|
Candidate
|
The accepted candidate. |
required |
journal
|
Any
|
The |
required |
speaker
|
Optional[str]
|
Attribution, passed through to the entry. |
None
|
topic_tags
|
Sequence[str]
|
Extra subject tags. The |
()
|
resolution
|
Optional[Resolution]
|
The M4 (#563) :class: |
None
|
Returns:
| Type | Description |
|---|---|
Optional[str]
|
The journal |
Optional[str]
|
one, else |
Optional[str]
|
failed -- in which case a terminal row records why). |
Source code in src/popoto/extraction/decision_log.py
607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 | |
get(agent_id, turn_id, candidate_id)
¶
Read one candidate's row, or None if it has none yet.
Source code in src/popoto/extraction/decision_log.py
list_for_agent(agent_id)
¶
Every row for an agent, in no particular order.
Reads by key pattern rather than through
DecisionRecord.query.filter(agent_id=...) on purpose, and
this is load-bearing rather than a style choice. Exact-match
KeyField filtering resolves through secondary index Sets that
KeyFieldMixin.on_save maintains
(fields/key_field_mixin.py:422+), and the guarded terminal
write is a deliberate ORM bypass -- it writes the row hash from
Lua, so on_save never runs and those index Sets are never
built. A candidate whose first write is terminal (every
firewall_drop / reject / withhold) would therefore be
invisible to filter() while sitting in Redis in full. Scanning
the keyspace has no such dependency.
Uses SCAN rather than KEYS so an unbounded log cannot block
the server. This is an operator/analysis reader, not a hot path.
Source code in src/popoto/extraction/decision_log.py
list_pending(agent_id, older_than=None)
¶
Stale pending rows for an agent, oldest-first.
The operator recovery reader (see the module docstring). A thin reader over existing rows -- it adds no keyspace, and decision rows carry no TTL, so nothing here deletes audit evidence.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_id
|
str
|
Owning agent. |
required |
older_than
|
Optional[float]
|
If given, only rows whose |
None
|
Returns:
| Type | Description |
|---|---|
List[DecisionRecord]
|
Matching rows sorted ascending by |
Source code in src/popoto/extraction/decision_log.py
turn_summary(agent_id, turn_id)
¶
The per-turn compact summary as {field: count}.
A convenience index over the detail rows, aggregating terminal states only. Correctness never depends on it -- if it ever disagrees with the detail rows, the detail rows are right.
Source code in src/popoto/extraction/decision_log.py
compute_metrics(agent_id, gold_labels)
¶
Precision/recall/F1 for one agent, from the log alone.
Reads decision-log rows and a caller-supplied gold-label mapping and nothing else -- no journal read, no LLM, no other live Redis state. That isolation is the point: it is what makes "extraction quality is computable offline" true rather than aspirational, and a test re-runs this with the journal keyspace flushed to prove it.
A row counts as a positive prediction when its state is accept.
Rows with no gold label are excluded from precision/recall (they are
still counted in the breakdowns), so a partially-labelled corpus
does not silently score as a pile of false positives.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
agent_id
|
str
|
Owning agent. |
required |
gold_labels
|
Dict[str, bool]
|
|
required |
Returns:
| Name | Type | Description |
|---|---|---|
A |
Metrics
|
class: |
Source code in src/popoto/extraction/decision_log.py
1038 1039 1040 1041 1042 1043 1044 1045 1046 1047 1048 1049 1050 1051 1052 1053 1054 1055 1056 1057 1058 1059 1060 1061 1062 1063 1064 1065 1066 1067 1068 1069 1070 1071 1072 1073 1074 1075 1076 1077 1078 1079 1080 1081 1082 1083 1084 1085 1086 1087 1088 1089 1090 1091 1092 1093 1094 1095 1096 1097 1098 1099 1100 1101 1102 1103 1104 1105 1106 1107 | |
Metrics
dataclass
¶
Offline extraction quality, computed from decision-log rows alone.
Attributes:
| Name | Type | Description |
|---|---|---|
precision |
float
|
|
recall |
float
|
|
f1 |
float
|
Harmonic mean of the two, 0.0 when undefined. |
true_positives |
int
|
Accepted and gold-labelled accept. |
false_positives |
int
|
Accepted and gold-labelled reject. |
false_negatives |
int
|
Not accepted but gold-labelled accept. |
per_reason_code |
Dict[str, int]
|
|
per_generator_rule |
Dict[str, Dict[str, int]]
|
|
Source code in src/popoto/extraction/decision_log.py
AuditableExtractionConfig
dataclass
¶
Opt-in configuration for the auditable extraction path.
Passing auditable_extraction=None (the default) is the only thing
the existing path ever sees, and its behavior is byte-for-byte
unchanged.
Deliberately carries no numeric knobs:
Defaults.M3_ASSEMBLY_CLAIM_TTL_MS is a pinned in-repo constant, not
a config field, per the repo's magic-number rule.
Attributes:
| Name | Type | Description |
|---|---|---|
verdict_provider |
Any
|
Anything callable as
|
journal |
Any
|
The :class: |
resolution_provider |
Any
|
Anything callable as
|
Source code in src/popoto/extraction/decision_log.py
hash_candidate_text(text)
¶
SHA-256 of the candidate text, for tamper-evidence on the row.