Risk scoring: how we learned to hear a polite threat
The first version looked for words like cancel and lawyer. It caught the obvious cases and missed the dangerous ones. Three versions later, risk is read from the shape of a conversation rather than its vocabulary.
By Howzer Team, Engineering ·
Consider two messages. The first is furious, in capitals, about a delivery being two days late. The second is calm, well written, and mentions in passing that this is the third time the writer has been in touch about the same thing. The first is loud. The second is the one about to leave. Getting that the right way round is the entire job of risk scoring.
Version one: a list of words
The first version mapped words to risk levels. Cancel, refund, lawyer. It was quick to build and it worked on messages that announce themselves. It could not tell the difference between mentioning a topic and intending it, so asking how cancellation works scored the same as announcing one.
Version two: words plus context
- Word signals combined with the sentiment of the message rather than read on their own.
- Repeated contact and a tone that worsens across messages counted as evidence.
- Thresholds calibrated separately for German and English.
- Far fewer false alarms on messages that were simply asking a question.
Version three: patterns
The current engine combines sentiment, emotion, root cause, the history of that customer and the structure of the message itself. The polite third contact triggers an escalation pattern despite its friendly tone, because the history weighs more than the wording.
What the level is for
- Someone about to leave: cancellation language, a competitor named, an ultimatum given.
- Legal and regulatory exposure: legal wording, a complaints authority mentioned, data protection invoked.
- Escalation building: contact after contact, each one a little sharper than the last.