Screenshots (OCR)
Read MoMo messages, and the sender, straight from screenshots. Needs pip install "cedikit[ocr]". Everything runs on your own computer: Windows' built-in OCR on Windows, RapidOCR elsewhere.
Tip
OCR can misread characters. cedikit fixes common slips (GHS50.OO becomes GHS50.00), but always compare the text with the picture.
cedikit.ocr
Read Mobile Money messages from screenshots (OCR), offline.
Needs pip install "cedikit[ocr]". On Windows it uses the OCR engine built into
Windows; elsewhere it uses RapidOCR. Both run entirely on this computer.
A screenshot of a messaging app holds more than the message: the clock, the
sender's name, "Today 10:04 AM", "Sender can't accept replies"... and often
several messages. :func:read_screenshot keeps only the message bubbles, splits
them into separate messages, fixes common OCR slips (GHS50.OO -> GHS50.00)
and picks out the sender shown at the top of the screen.
Example::
from cedikit import ocr, fraud
shot = ocr.read_screenshot("screenshot.png")
for message in shot.messages:
print(fraud.check(message, sender=shot.sender).risk)
OCR can misread text: always compare the result with the picture.
OcrUnavailable
Bases: CedikitError
Raised when no OCR engine is installed.
OcrLine
dataclass
One line of text found in an image, with its position in pixels.
Screenshot
dataclass
What was read from a screenshot.
Attributes:
| Name | Type | Description |
|---|---|---|
messages |
list[str]
|
The message bubbles, top to bottom, as text (OCR slips fixed). |
sender |
str | None
|
The sender shown at the top of the screen, if one was recognised
(an official sender ID such as |
engine |
str
|
Which OCR engine was used. |
lines |
list[OcrLine]
|
Every line the engine found (for troubleshooting). |
fix_ocr_text(text)
Undo common OCR confusions in MoMo messages, mainly letters read in numbers.
Only money amounts and a few fixed phrases are touched, so names and other words are left exactly as read.
Example
fix_ocr_text("Cash Out made for GHS50.OO. Financial Transaction ld: 191") 'Cash Out made for GHS50.00. Financial Transaction Id: 191'
group_messages(lines, image_height)
Split OCR lines into message bubbles and find the sender.
Lines in the header (top of the screen), phone chrome such as clocks and "Sender can't accept replies", and stray icons are dropped. A vertical gap wider than about 1.3 line-heights starts a new message.
Returns:
| Type | Description |
|---|---|
tuple[list[str], str | None]
|
|
available_engine()
The OCR engine that will be used, or None if none is installed.
read_screenshot(image)
Read the MoMo messages in a screenshot (a file path or the image bytes).
Raises:
| Type | Description |
|---|---|
OcrUnavailable
|
If no OCR engine or image library is installed. |