STUDY / 01 · Discover Data 4:20, 2026
TEMPLAR-Fraud
A verifier-grounded, calibrated and cost-sensitive framework for transaction fraud detection. This page separates what the paper reports from what this account's reimplementation measures on synthetic data.
Contribution: baseline comparison, evaluation-metric analysis, robustness interpretation, and manuscript refinement.
Interactive
Architecture explorer
Select a stage to see its named inputs, outputs and source section. When the repository's demo trace is available the explorer shows the recorded step for a sampled transaction; otherwise it is labelled schematic.
Section 2.2, Table 2
Input fields
Preprocessing is fitted on the training fold only: median imputation, an explicit missing category, one-hot for low-cardinality and frequency encoding for high-cardinality columns, robust scaling. Temporal features are a month index and day of week where the data has it.
Inputs
- timestamp
- amount
- categorical descriptors
- late fraud label (training only)
Outputs
- feature vector after train-fold imputation, encoding and scaling
Dataset statements
- BAF is a synthetic bank account-opening benchmark (Bank Account Fraud, NeurIPS 2022 Datasets and Benchmarks). It is not real customer data and does not contain card, device or IP columns; its only temporal field is a month index.
- IEEE-CIS is the Kaggle IEEE-CIS Fraud Detection dataset, obtained through Kaggle under its competition terms.
- Neither dataset is redistributed by the repository. Real-data adapters validate a schema and raise a clear error naming what to obtain and from where.
Protocol notes
- Chronological 70/15/15 split; forward-chaining validation for model selection; preprocessing fitted on the training fold only.
- The late slice is the final 10% of rows in time and sits inside the 15% test partition, so late-slice metrics are not independent of test metrics. The repository measures this overlap rather than assuming it.
- Results are means over five seeds. Hyperparameters are selected on validation average precision.
- The decision threshold is chosen on validation to minimise a cost with three weights; the paper does not state the units of those weights.
Paper versus this account's reimplementation
The metrics in the table are paper-reported and were produced by the paper's authors on the benchmarks named. The repository linked above is an independent implementation written from the paper's description. It runs end to end on an authored synthetic fixture, and every number it produces carries the label demo. Nothing on this site is labelled reproduced, because no benchmark data has been processed here.
Ambiguities logged by the reimplementation
- The paper's late slice is the final 10% of the same dataset and sits inside the 15% test partition, so it measures within-dataset drift rather than transfer to another population.
- Cost units for the decision rule are not stated in the paper; the lab uses stated demo cost weights that can be changed but are not the paper's.
- BAF carries a month index only, so the day-of-week and card, device and IP graph features described for IEEE-CIS have no equivalent there.
- The template bank in the reimplementation is hand-written; the paper describes a search-based discovery procedure that is not reproduced.
- Sample counts for the calibration fit and the exact routing thresholds are reported inconsistently in the source and are logged as ambiguities in the repository's fidelity audit.