Evaluation
Measure parser accuracy and fraud precision/recall on labelled files. Always evaluate on a held-out set that was not used to write rules or train models.
cedikit.evaluation
Measure the parser and fraud checker against labelled, anonymised messages.
The file format is the one used by tests/fixtures/sample_messages: a YAML
list of entries with text, sender and (for parser checks) expected
fields. Keep a held-out set that is never used to tune rules or train models.
Example::
from cedikit import evaluation
genuine = evaluation.load("genuine.yaml")
scam = evaluation.load("scam.yaml")
print(evaluation.evaluate_fraud(genuine, scam))
print(evaluation.evaluate_parser(genuine))
FraudEvaluation
dataclass
recall
property
Share of scams that were flagged.
precision
property
Share of flagged messages that really were scams.
load(path)
Load a labelled YAML file.
evaluate_fraud(genuine, scam, *, flag_at='MEDIUM', use_sender=True, classifier=None)
Count how many scams are flagged and how many genuine messages are wrongly flagged.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
flag_at
|
Risk
|
The lowest risk level counted as "flagged". |
'MEDIUM'
|
use_sender
|
bool
|
Pass each message's sender to the checker. Set False to measure how well the text alone is judged. |
True
|
field_mismatches(tx, expected)
Fields of tx that differ from expected, as readable strings.
evaluate_parser(samples)
Check that each sample parses with all its expected fields correct.