# Payment data files: method and field guide

Version 1.1 · 2026-10-01 · Simpa Labs

All rows are invented. These files contain no customer data and measure no live product. They are hand-authored fixtures, so there is no random seed. CSV files use UTF-8, comma separators, and a header. Empty cells mean not supplied or not applicable; they never mean zero. Copy fixtures only into a test environment. Use a new file version when changing a field's meaning. Keep the version used with each test result.

## Synthetic Nigerian payments

`synthetic-nigerian-payments-v1.csv` contains 16 rows and 13 operations. The filename remains stable; `dataset_version` identifies revision 1.1. Existing fields remain available. The test model records a single intended money effect; it does not describe every line of a double-entry journal.

| Field | Meaning |
| --- | --- |
| record_id | Unique fixture row ID |
| operation_id | One intended business action; shared on retry |
| provider_event_id | Provider message ID; blank when no message was sent |
| amount_ngn | Intended whole-naira amount; positive integer |
| channel | Test path; a fixture label, not a provider's API enum |
| status | Scenario input label, not a universal payment state |
| expected_outcome | Action to check in the tested application |
| source | Always synthetic |
| dataset_version | Schema and fixture revision |
| scenario_kind | Local request, lost response, success, retry, or redelivery |
| observed_amount_ngn | Amount in the input; compare with intended amount |
| currency | NGN |
| prior_effect_count | Money effects already committed before this row |
| expected_total_effect_count | Total effects after handling this row |
| provider_state | not_sent, unknown, or provider_success |
| expected_state | confirmed, rejected, unresolved, or review |

Apply each row with its stated precondition. Do not add expected totals across rows. For `op_a`, row 001 starts with zero effects and ends with one; row 002 starts with one and must keep one. The mismatch row has intended NGN 8000 and observed NGN 7500; hold it for review. Local rejection rows have no provider event ID. Timeout rows remain unresolved with zero effects at this checkpoint; a later confirmed success needs a new fixture. The loan row treats NGN 2000 as a partial repayment against an NGN 2000 request; contract allocation must be supplied by your test setup. This file does not contain balances, loan terms, device records, account assignment history, or signed callbacks. Add those preconditions before testing the full rule.

## Duplicate risk inputs

`duplicate-risk-synthetic.csv` contains 10 event rows for five intended operations. `event_id` identifies a delivery or request, while `provider_event_id` identifies the underlying provider message. The last two rows have different delivery IDs and the same provider message ID. That is a redelivery. The first redelivery fixture presumes an earlier provider success message outside this small input set; supply it in your adapter setup.

`event_type` names the injection point. `amount_ngn` is the intended whole-naira amount. `event_time_utc` orders the input sequence. `expected_effects` is the expected final operation count, repeated on each input row for ease of checking; never sum it. `scenario` describes the fixture. `source` is synthetic. Keep results in a separate table with one row per operation: run_id, operation_id, build, observed_effects, wait_window_seconds, evidence_uri.

For these fixtures the expected final count is one for each of five operations. Count a business effect, not both balancing journal lines. A credit and its balancing debit form one intended posting. Approved reversals and distinct partial refunds are separate business actions with their own operation IDs.

Duplicate-effect rate = operations with observed effects above their expected effects / completed operations tested. Extra-effect count = sum of max(observed effects - expected effects, 0). Missing-effect count is separate. If one of five operations has two effects and four have one, the fixture rate is 1/5 = 20% and the extra-effect count is one. This is a hypothetical test result, never a market estimate. State the wait window and leave unresolved operations outside the completed denominator; report their count beside it.

## Incident timeline

`payment-incident-timeline-template.csv` has eight fictional rows plus a blank row. Delete all fictional rows for real use. `row_kind` labels example data. `case_id` groups the case. `event_time_utc` is source event time; `observed_time_utc` is receipt time. Both use RFC 3339 UTC strings ending in Z. `original_time_zone` records the source zone. `clock_offset_ms` is source-clock minus reference-clock offset; blank means unknown. Do not adjust the fictional zero offsets.

`source` and `event_id` identify a record. `operation_id`, `provider_reference`, and `ledger_entry_id` join records when available. `status` is the source state. `confirmed_fact` states only evidence-backed facts. `hypothesis` stores an explanation to test; `none` means none recorded. `evidence_uri` points to a restricted retained record. The example:// links are placeholders and do not resolve. `customer_effect` and `money_position` show impact and reconciliation state. `owner` and `action` name the next step.

The fictional sequence goes from accepted request to delayed ledger credit and closure. The provider redelivery retains its original event time and has a later observed time. Sort by observed time to review response work and by event time to review source order. Never infer cause from timestamps alone. Preserve raw timestamps and known clock error with the source evidence.

## Evidence index

`fintech-evidence-index-v1.csv` has six example controls. Every status is example_only, every evidence link is blank, and no control is claimed to pass. `index_version` is the format revision. `control_id` stays stable; `control_version` changes when the rule changes. `control_statement` is one testable rule. `owner_role` names who maintains it. `source_reference` is an actual reference URL, not proof of compliance.

`evidence_type` describes the proposed test. Use separate `design_evidence_uri`, `test_evidence_uri`, `run_evidence_uri`, and `exception_evidence_uri` links. An exception link is optional when no exception exists. `evidence_location` is a legacy general location field, retained for compatibility. `system_build` identifies the tested release; `sample_scope` names records and review period. `review_date` and `valid_until` use YYYY-MM-DD. `review_trigger` states when evidence must be refreshed. `reviewer_role` names who checked it. `retention_rule` points to your retention policy.

Use status current only after an authorized reviewer opens the required evidence, checks scope and build, and sets review dates. Use missing for absent required records, expired for stale evidence, failed for a recorded failed test, and not_applicable only with a reason. Fill `status_reason` for every status. Expired evidence cannot be made current by changing a date alone. Publicly share the index only after removing real customer data, private links, and secrets.

## License and changes

These four CSV files and this method file are dedicated under CC0 1.0: https://creativecommons.org/publicdomain/zero/1.0/. Attribution is welcome but not required. The related source pages have their own terms.

Revision 1.1 adds explicit preconditions, effect counts, provider message identity, chronological incident examples, and separate evidence slots. Existing CSV filenames and original columns remain. New columns require adapters to read fields by header rather than column position. Preserve a copy of each downloaded revision before editing it.

## Reference basis

- Stripe webhook guidance describes duplicate deliveries and event ordering: https://docs.stripe.com/webhooks
- OWASP logging guidance describes event fields and handling sensitive data: https://cheatsheetseries.owasp.org/cheatsheets/Logging_Cheat_Sheet.html

The fixture rules and example counts are supplied by this method, not measured from those sources.
