Md Sultanul Arefin Sourav

STUDY / 01 · Demo ready

TEMPLAR-Fraud lab

Independent implementation of a routed, verifier-grounded, calibrated and cost-sensitive fraud decision pipeline with synthetic demo data and a paper fidelity audit.

researchdemo readyfraud-detectioncalibrationcost-sensitive-learningexplainabilitypython

What decision is supported

Whether a transaction or an account-opening application is approved automatically, held for human review, or declined, and at which score threshold the expected cost of wrong decisions is lowest.

Who uses it

A fraud operations analyst who owns the review queue, and a model owner who has to justify the operating threshold to an audit function.

What data enters

Tabular records with a timestamp, an amount, categorical descriptors and a fraud label that arrives late. The published study used BAF, a synthetic account-opening benchmark, and IEEE-CIS via Kaggle. This repository ships only an authored synthetic fixture; no benchmark rows are redistributed.

What is computed

A triage probability from a gradient-boosted model, an uncertainty estimate, a route (fast path or escalation), a verifier decision on any generated rationale, a stacked probability over the available sources, a calibrated probability, and a decision threshold chosen to minimise a stated cost.

What action is suggested

Approve, review or decline. A rationale is shown only if the verifier accepted it; a rejected rationale sends the case to the human review queue instead of being displayed.

What evidence supports it

Evidence for TEMPLAR-Fraud lab
MetricDatasetSliceValueLabelSource
AUROCBAFinternal chronological test0.918paper-reportedAbstract; Tables 3 and 4
AUC-PRBAFinternal chronological test0.498paper-reportedAbstract; Tables 3 and 4
F1BAFinternal chronological test0.557paper-reportedAbstract; Tables 3 and 4
AUROCIEEE-CISinternal chronological test0.972paper-reportedAbstract; Tables 3 and 4
AUC-PRIEEE-CISinternal chronological test0.701paper-reportedAbstract; Tables 3 and 4
F1IEEE-CISinternal chronological test0.747paper-reportedAbstract; Tables 3 and 4
AUROCBAFlate 10% chronological slice0.901paper-reportedTables 7 and 8
AUC-PRBAFlate 10% chronological slice0.452paper-reportedTables 7 and 8
F1BAFlate 10% chronological slice0.518paper-reportedTables 7 and 8
AUROCIEEE-CISlate 10% chronological slice0.951paper-reportedTables 7 and 8
AUC-PRIEEE-CISlate 10% chronological slice0.642paper-reportedTables 7 and 8
F1IEEE-CISlate 10% chronological slice0.687paper-reportedTables 7 and 8
ECE after calibrationBAFinternal chronological test0.014paper-reportedTables 12 and 13
ECE after calibrationIEEE-CISinternal chronological test0.013paper-reportedTables 12 and 13

Every value above is reported by the paper and was not produced by this account's code. The reimplementation runs end to end on a synthetic fixture and its outputs carry the label demo; nothing in it is labelled reproduced because no benchmark data was processed.

What fails or is uncertain

  • The paper's late slice is the final 10% of the same dataset and sits inside the 15% test partition, so it measures within-dataset drift rather than transfer to another population.
  • Cost units for the decision rule are not stated in the paper; the lab uses stated demo cost weights that can be changed but are not the paper's.
  • BAF carries a month index only, so the day-of-week and card, device and IP graph features described for IEEE-CIS have no equivalent there.
  • The template bank in the reimplementation is hand-written; the paper describes a search-based discovery procedure that is not reproduced.
  • Sample counts for the calibration fit and the exact routing thresholds are reported inconsistently in the source and are logged as ambiguities in the repository's fidelity audit.