DEMO / 08 · Demo ready
Complaint intelligence lab
Complaint categorisation and review routing on the CFPB schema with a TF-IDF baseline, abstention thresholds and chronology-aware evaluation.
What decision is supported
Which team a consumer complaint should be routed to, and when the classifier should abstain and hand the text to a person instead of guessing.
Who uses it
A complaints operations lead at a fictional financial services firm, and an analyst who wants to know how routing accuracy changes over time rather than on a shuffled split.
What data enters
Text and metadata in the shape of the CFPB consumer complaint schema. The repository ships authored synthetic complaints only; the public CFPB data can be fetched by the user under the provider's terms.
What is computed
A TF-IDF and linear classifier baseline, a confidence threshold below which the system abstains, an evaluation that trains on earlier months and tests on later months, and theme summaries per category from the fitted vocabulary.
What action is suggested
A routing label with its confidence, or an explicit abstention that places the complaint in a human review queue, plus a per-period report of coverage and accuracy at the chosen threshold.
What evidence supports it
Coverage and accuracy trade-offs are shown for the synthetic fixture only, labelled demo.
No metric is claimed for this project. Anything computed by the repository on its synthetic fixture carries the label demo.
What fails or is uncertain
- Synthetic complaint text is templated, so vocabulary is narrower than real complaints.
- The baseline is linear; the lab does not claim any result for transformer models.
- Themes are the top weighted terms of a linear model, not a topic model.