| Download | - View final version: OCR evaluation tools for the 21st century (PDF, 238 KiB)
|
|---|
| Author | Search for: Santos, Eddie Antonio1 |
|---|
| Affiliation | - National Research Council Canada. Digital Technologies
|
|---|
| Format | Text, Article |
|---|
| Conference | 3rd Workshop on the Use of Computational Methods in the Study of Endangered Languages, February 26-27, 2019, Honolulu, Hawaii |
|---|
| Abstract | We introduce ocreval, a port of the ISRI OCR Evaluation Tools, now with Unicode support. We describe how we upgraded the ISRI OCR Evaluation Tools to support modern text processing tasks. ocreval supports producing character-level and word-level accuracy reports, supporting all characters representable in the UTF-8 character encoding scheme. In addition, we have implemented the Unicode default word boundary specification in order to support word-level accuracy reports for a broad range of writing systems. We argue that character-level and word-level accuracy reports produce confusion matrices that are useful for tasks beyond OCR evaluation— including tasks supporting the study and computational modeling of endangered languages. |
|---|
| Publication date | 2019-02 |
|---|
| Publisher | Association for Computational Linguistics |
|---|
| Licence | |
|---|
| In | |
|---|
| Language | English |
|---|
| Peer reviewed | Yes |
|---|
| Export citation | Export as RIS |
|---|
| Report a correction | Report a correction (opens in a new tab) |
|---|
| Record identifier | 9ed97177-7f1d-4955-b2c3-b11bd4416187 |
|---|
| Record created | 2022-07-29 |
|---|
| Record modified | 2022-07-29 |
|---|