we_are_coded.by CODE · The world, decoded
БГ
Concept

OCR: reading text off a photo

The BasicsUpdated on 16 August 2026we are coded

Paper kept its secrets long enough. Now AI makes it talk.

Checked on16 August 2026
In short: OCR (Optical Character Recognition) is the technology that reads letters off a photo or scanned page and turns them into real text you can search, copy and edit. It has existed for decades, but until recently it only handled clean, typewritten text on a straight page well. AI-powered models made it sharply better, because they no longer recognize letters one by one - they read the whole page as meaning. That's why an old contract, a handwritten note or a faded invoice can suddenly talk.

The technology starts from one very specific, human need. In 1976, American inventor Ray Kurzweil unveiled a machine that scanned a printed page and read it aloud - so blind people could "read" ordinary books. That's the first practical use of OCR, period. Since then the technology has passed through banks, libraries, court archives, but for decades stayed clumsy and temperamental.

Classic OCR works as a strict shape matcher. It learns to compare every letter to a template, and when the shape matches closely enough, it records it. It works decently on a cleanly scanned page in a familiar font. A crooked phone photo, a handwritten signature, a coffee stain on an invoice or a table with crooked columns - and it breaks almost instantly. The result is a string of wrong characters that make the whole text meaningless.

AI flips the logic. Instead of recognizing isolated shapes, the model is trained on a huge number of text images and learns to guess what's written from context - exactly like large language models, LLMs, systems trained to predict the next word in a sentence, guess the next word in a conversation. Is the letter blurry, but the sentence obviously continues with a specific word? The model recognizes it by meaning, not just shape. That's why today's OCR reads handwriting, crumpled receipts, photos in bad light and even pages with tables and stamps that used to make older systems give up.

This is where the technology meets its real job: all the paper in the world. Contracts in filing cabinets, medical records, old letters, municipal archives, accounting documents - all of it can suddenly turn into a searchable digital archive, instead of sitting locked in a folder no one will ever open again. Free, open-source tools like Tesseract have done the basic, rough work for years, but AI models add the understanding that was missing.

There's also a cost to this ease. A model that "guesses" meaning instead of reading literally sometimes guesses wrong - and sounds fully confident in its mistake. For a photo of a note, that's a minor risk. For a medical referral, a court document or a contract with numbers, it's a different question. The error looks like the truth until someone checks the original.

The visual is generated code art. No third-party images.
Official primary sources
→Google Cloud: Vision OCR - documentation