Skip to main content Skip to main navigation menu Skip to site footer
##common.pageHeaderLogo.altText##
Izvestiya SFedU
Engineering sciences
  • Current
  • Previous issues
    • Archive
    • Issues 1995 – 2019
  • Editorial Board
  • About journal
    • Officially
    • The main tasks
    • Main sections
    • Specialties of the Higher Attestation Commission of the Russian Federation
    • Editor-in-Chief
ISSN 1999-9429 print
ISSN 2311-3103 online
  • Login
  1. Home /
  2. Search

Search

Advanced filters
Published After
Published Before

Search Results

##search.searchResults.foundPlural##
  • ON THE ACCURACY AND COMPLEXITY OF THE MULTI-STAGE METHOD FOR CORRECTING DISTORTED TEXTS DEPENDING ON THE DEGREE OF DISTORTION

    D.V. Vakhlakov, V. А. Peresypkin, А.V. Germanovich, S.Y. Melnikov, N.N. Copkalo
    130-142
    2021-10-05
    Abstract ▼

    One of the main factors that significantly complicate the understanding, translation and analysis of texts obtained by automatic recognition of speech or images of texts is the presence of distortions in the form of erroneous symbols, words and phrases. Until recently, there were no effective software tools for correcting texts with significant distortions, although this task is rele-vant both for Russian and other common languages in the context of the active use of recognition systems in advanced augmented reality systems. The authors proposed a new multi-stage method for correcting distorted texts, which significantly increases the accuracy of the correction (in terms of the number of correctly corrected words in the text) and is based on the sequential detec-tion of errors and their correction. In this paper, we evaluate the accuracy and computational complexity of the proposed method for correcting distorted texts at various levels of distortion, and determine its place among other modern approaches to correction. The most typical errors of recognition systems are: – replacing a word with a similar sound or graphic spelling; – replacing several words with one; – replacing one word with several; – omission of words; – insertion or deletion of short words (including prepositions and conjunctions). As a result of recognition, a distorted text is obtained, which consists mainly of dictionary words, even in places of distortion. With a large number of distortions, the texts become almost unreadable. Due to the fact that it is problematic to select texts with a wide range of distortion levels in the required amount based on the results of real machine recognition of speech and images of texts, software modeling of distor-tions was used. A text distortion technique has been proposed and implemented that simulates the results of recognition systems in a wide range of distortions; distorted texts have been prepared in the required amount. Within the framework of the proposed multi-stage correction method, non-dictionary word forms and words are considered distorted if the probability of their occurrence in the text in accordance with the chosen language model is less than a given threshold. For such distorted words, a list of possible variants of words is built, which includes only those word forms from the dictionary that are at a certain Levenshtein distance from the word under study. The cor-rected text from the tables of word variants is obtained by searching for the most probable chain of word forms. The correction method consists of several stages, at each stage only those frag-ments of the text that remain distorted after the previous stage are corrected. According to the results of the experiments on the correction of distorted texts, it was concluded that the proposed correction method showed good results with an average value of F-measure >50 % in the distor-tion range from 0 to 75 %. Linguistic experts confirmed the fruitfulness of the proposed approach to correction and its preference over other modern approaches, fixing that with a level of distor-tion of up to 50 % of words, the corrected text is read with much less effort than a distorted one, and with a level of distortion of up to 70% of words, the corrected text also allows you to highlight useful information about the content

  • MULTI-PASS METHOD FOR AUTOMATIC CORRECTION OF DISTORTED TEXTS

    D.V. Vakhlakov, V.A. Peresypkin, S.Y. Melnikov
    2021-02-25
    Abstract ▼

    One of the main factors that significantly complicate the understanding, translation and
    analysis of texts obtained by automatic speech recognition or optical recognition of text images
    are the distortions contained in them in the form of erroneous characters, words and phrases.
    The most typical errors of recognition systems are: – replacement of a word with a similar sounding
    or graphic spelling; – replacing several words with one; – replacement of one word with several;
    – skipping words; – insertion or deletion of short words (including prepositions and conjunctions).
    As a result of recognition, a text is obtained that has distortions and consists mainly of dictionary
    words, including in places of distortion. With a large amount of distortion, the texts become
    almost unreadable. Automatic processing of such texts is very difficult, although this task is
    relevant both for Russian and for other common languages. Correction software that works well at
    low distortions in the text, in the case of texts with a high level of distortion, regardless of their
    origin, show unsatisfactory results. This makes it necessary to develop independent approaches to
    correcting distorted texts. A new multi-pass method for correction of distorted texts based on sequential
    error identification and correction of distorted texts is proposed. Non-dictionary word
    forms and word forms which occurrence probability in the text in accordance with the selected
    probabilistic model is less than a preset threshold are considered to be distorted. After setting of
    the distortion sign for individual words, this sign is spread to their combinations, i.e. distorted text
    fragments are extracted. A list of possible word variants which includes only those word forms
    from the dictionary that are located at a certain Levenshtein distance from the word under study is
    built for them. The corrected text from word variants is obtained by searching for the most probable
    chain of word forms. The correction method consists of several passes, at each pass only those
    fragments of the text are corrected that remained distorted after the previous pass of correction.
    The method allows to increase significantly the quality (accuracy) of the correction. In the carried
    out experiments the quality of correction in terms of the F1-measure for moderately distorted texts
    has been increased by 9 %, and for highly distorted texts – by 7.7 %.

1 - 2 of 2 items

links

For authors
  • Submit article
  • Author Guidelines
  • Editorial Policy
  • Reviewing
  • Ethics of scientific publications
  • Open access policy
  • Supporting documents
Language
  • English
  • русский

journal

* not an advertisement

index

Индексация журнала
* not an advertisement
Information
  • For Readers
  • For Authors
  • For Librarians
Address: 347900, Taganrog, Chekhov St., 22, A-211 Phone: +7 (8634) 37-19-80 E-mail: iborodyanskiy@sfedu.ru
Publication is free
More information about the publishing system, Platform and Workflow by OJS/PKP.
logo Developed by RDCenter