AIThe Decoder1h ago

Old OCR text cripples language model training, and FineBooks wants

Old OCR text cripples language model training, and FineBooks wants to fix that at scale

Old OCR text cripples language model training, and FineBooks wants

The FineBooks project from Hugging Face and EleutherAI tested 14 open-source OCR models on more than 2,000 historical book pages. The top model, dots.mocr, hits 97.6 percent character accuracy at under two dollars per thousand pages. That's good enough for AI training data, but…

Read full article

Source: The Decoder · Opens in new tab