Skip to content

Evaluation

Measure parser accuracy and fraud precision/recall on labelled files. Always evaluate on a held-out set that was not used to write rules or train models.

cedikit.evaluation

Measure the parser and fraud checker against labelled, anonymised messages.

The file format is the one used by tests/fixtures/sample_messages: a YAML list of entries with text, sender and (for parser checks) expected fields. Keep a held-out set that is never used to tune rules or train models.

Example::

from cedikit import evaluation
genuine = evaluation.load("genuine.yaml")
scam = evaluation.load("scam.yaml")
print(evaluation.evaluate_fraud(genuine, scam))
print(evaluation.evaluate_parser(genuine))

FraudEvaluation dataclass

recall property

Share of scams that were flagged.

precision property

Share of flagged messages that really were scams.

load(path)

Load a labelled YAML file.

evaluate_fraud(genuine, scam, *, flag_at='MEDIUM', use_sender=True, classifier=None)

Count how many scams are flagged and how many genuine messages are wrongly flagged.

Parameters:

Name Type Description Default
flag_at Risk

The lowest risk level counted as "flagged".

'MEDIUM'
use_sender bool

Pass each message's sender to the checker. Set False to measure how well the text alone is judged.

True

field_mismatches(tx, expected)

Fields of tx that differ from expected, as readable strings.

evaluate_parser(samples)

Check that each sample parses with all its expected fields correct.