STUDY / 01 · Demo ready
TEMPLAR-Fraud lab
Independent implementation of a routed, verifier-grounded, calibrated and cost-sensitive fraud decision pipeline with synthetic demo data and a paper fidelity audit.
What decision is supported
Whether a transaction or an account-opening application is approved automatically, held for human review, or declined, and at which score threshold the expected cost of wrong decisions is lowest.
Who uses it
A fraud operations analyst who owns the review queue, and a model owner who has to justify the operating threshold to an audit function.
What data enters
Tabular records with a timestamp, an amount, categorical descriptors and a fraud label that arrives late. The published study used BAF, a synthetic account-opening benchmark, and IEEE-CIS via Kaggle. This repository ships only an authored synthetic fixture; no benchmark rows are redistributed.
What is computed
A triage probability from a gradient-boosted model, an uncertainty estimate, a route (fast path or escalation), a verifier decision on any generated rationale, a stacked probability over the available sources, a calibrated probability, and a decision threshold chosen to minimise a stated cost.
What action is suggested
Approve, review or decline. A rationale is shown only if the verifier accepted it; a rejected rationale sends the case to the human review queue instead of being displayed.
What evidence supports it
| Metric | Dataset | Slice | Value | Label | Source |
|---|---|---|---|---|---|
| AUROC | BAF | internal chronological test | 0.918 | paper-reported | Abstract; Tables 3 and 4 |
| AUC-PR | BAF | internal chronological test | 0.498 | paper-reported | Abstract; Tables 3 and 4 |
| F1 | BAF | internal chronological test | 0.557 | paper-reported | Abstract; Tables 3 and 4 |
| AUROC | IEEE-CIS | internal chronological test | 0.972 | paper-reported | Abstract; Tables 3 and 4 |
| AUC-PR | IEEE-CIS | internal chronological test | 0.701 | paper-reported | Abstract; Tables 3 and 4 |
| F1 | IEEE-CIS | internal chronological test | 0.747 | paper-reported | Abstract; Tables 3 and 4 |
| AUROC | BAF | late 10% chronological slice | 0.901 | paper-reported | Tables 7 and 8 |
| AUC-PR | BAF | late 10% chronological slice | 0.452 | paper-reported | Tables 7 and 8 |
| F1 | BAF | late 10% chronological slice | 0.518 | paper-reported | Tables 7 and 8 |
| AUROC | IEEE-CIS | late 10% chronological slice | 0.951 | paper-reported | Tables 7 and 8 |
| AUC-PR | IEEE-CIS | late 10% chronological slice | 0.642 | paper-reported | Tables 7 and 8 |
| F1 | IEEE-CIS | late 10% chronological slice | 0.687 | paper-reported | Tables 7 and 8 |
| ECE after calibration | BAF | internal chronological test | 0.014 | paper-reported | Tables 12 and 13 |
| ECE after calibration | IEEE-CIS | internal chronological test | 0.013 | paper-reported | Tables 12 and 13 |
Every value above is reported by the paper and was not produced by this account's code. The reimplementation runs end to end on a synthetic fixture and its outputs carry the label demo; nothing in it is labelled reproduced because no benchmark data was processed.
What fails or is uncertain
- The paper's late slice is the final 10% of the same dataset and sits inside the 15% test partition, so it measures within-dataset drift rather than transfer to another population.
- Cost units for the decision rule are not stated in the paper; the lab uses stated demo cost weights that can be changed but are not the paper's.
- BAF carries a month index only, so the day-of-week and card, device and IP graph features described for IEEE-CIS have no equivalent there.
- The template bank in the reimplementation is hand-written; the paper describes a search-based discovery procedure that is not reproduced.
- Sample counts for the calibration fit and the exact routing thresholds are reported inconsistently in the source and are logged as ambiguities in the repository's fidelity audit.