检索规则说明:AND代表“并且”;OR代表“或者”;NOT代表“不包含”;(注意必须大写,运算符两边需空一格)
检 索 范 例 :范例一: (K=图书馆学 OR K=情报学) AND A=范并思 范例二:J=计算机应用与软件 AND (U=C++ OR U=Basic) NOT M=Visual
机构地区:[1]广西计算中心,广西南宁530022
出 处:《广西科学院学报》2011年第4期317-319,共3页Journal of Guangxi Academy of Sciences
摘 要:设计一种基于多引擎的印刷体汉字识别系统,优先采用汉王光学字符识别(OCR)引擎的版面分析结果,在汉王、清华OCR引擎分别完成字符识别之后,根据字符的图像坐标,整合两者的识别结果,并用彩色突出两OCR引擎的冲突字符、置信度低的字符及WiseCheck语义校对引擎提示的错误字符。该系统改善了现有大规模数字化加工生产线中人工比照图像时对识别文本逐字、全文遍历式校对的工作模式,能减轻劳动强度,提高工作效率,降低处理成本。A printed Chinese characters recognition system based on multi-engine has been constructed.Basing on the HW-OCR engine's layout analysis,the HW-OCR and TH-OCR engines accomplished character recognition respectively.According to the coordinate of the character image,the system will integrate the two OCR engine's recognition results using different colors to highlight their conflict character and low confidence character,and the other wrong words which are checked by the "WiseCheck"(a semantic collation engine).This system has improved the text verbatim identification by artificial contrast image and full-text search proofreading work mode in the existing mass digitization processing production line,which further can reduce labor intensity,improve work efficiency and reduce the cost of processing.
分 类 号:TP391.1[自动化与计算机技术—计算机应用技术]
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在载入数据...
正在链接到云南高校图书馆文献保障联盟下载...
云南高校图书馆联盟文献共享服务平台 版权所有©
您的IP:216.73.216.233