MAPS (Onomastics)

Synopsis

MAPSOno is the family name of our anthroponomy processing package; a set of specialized modules tuned for applications such as information retrieval, document clustering, named entity extraction (NER), and translation.

MAPS Toponymy

MAPS Orthography

MAPS Semantics

Information

Last updated: 7/4/2024

MAPS Onomastics

Romanization is a way of reproducing the sound of the words according to the orthography rules of the target language; uniform results in the Romanization of names are difficult to obtain, since vowel points and diacritical marks are generally omitted from both manual and machine writing in some languages like Arabic. It follows that for correct identification of the words which appear in any particular name, knowledge of its standard Arabic-script spelling including proper pointing, and recognition of dialectal and idiosyncratic deviations are essential. Please refer Personal Names Romanizer for details.

Names romanizer

Representing personal names written in different languages and vice versa is a task always been described as a challenge to most cross-language content management and data mining systems; MAPS works not only for Romanization but also for a dozen of languages, the integral transliteration system makes it possible to take names in native script or Romanized form, perform the transcription and return results in the target native script, the output "transcribed names" are formatted in a special way directed to readers of the particular geographic region; For instance the Arabic name "بُرْهَان" is rendered "Бурхан" for Russian, "Burhan" for English speakers, Czech or Spanish, "Borhane" for Francophone and "Borhan" for German and Polish. Please refer to the Personal Names Transcriber for details and samples.

Names transcriber

Like Romanization, the reverse way "Arabicization" is also possible with Kalmasoft's Names Arabicization System, a module that produces the vocalized Arabic version of non-Arab names written in any of the European Union's languages as well as many other Semitic and Afro-Semitic languages in their native script form.
As the adoption of a non-Arab name into one language is usually a process of adjusting its original pronunciation to suit the phonological regularities in the target language, the name undergoes significant transformation to accommodate the phonological characteristics of the target language.

Distinction here should be made between two important points:

  • Arabic phonetic system accepts foreign imported phonemes and represents them in a way good enough for trained reader and native speakers to spell the name and pronounce it close to the original except for some phonemes and digraphs e.g. (P,V) and ("CH",""); that is true regarding the extra set of short vowels our system is making the best use of.
    Unlike Arabic Romanization, variants here occur as a result of Arabic varieties used in different geographic regions; our system adopts the MSA in general but can generate other varieties of Arabic as well.
  • The second point can be clarified using the name Michael which is actually spelled "مايكل" in Arabic but as "ميشيل" and "ميشال" too, the later is common in Levantine, similar example is the name Nichola which yields "نيكولا" and "نقولا"; other example is the name Clinton" is commonly spelled as "كلنتون" which will work just fine but another variant does exist "كلنطون" which is common in north Africa, Egypt in particular; the samples in the link below will show some other nuances. MAPS deals not only with these but also with other complexities between different varieties of Arabic, please refer to Personal Names Arabicizer for details.

Names Arabicizer

A reversible name transcription system is not easy to accomplish, the name's inherent orthographic characteristics are totally lost when exported to another language. This is a major challenge for information retrieval systems utilized in CLIR applications.
No matter how bad the changes distorted the original orthography our system will reproduce the original name thus improving the accuracy of name searching, transliteration, and the quality of identity verification initiatives by providing high accuracy ranked results, based on the linguistic, phonetic, and specific cultural variation patterns of the names. Please refer to Personal Names Retriever for detailed output sample.

Names retriever

Currently applicable to Arabic names only, this module makes use of the truth that different geographic regions have different name patterns and most have specific set of names unique to it beside other patterns that are in common e.g. the names "حفني" /ħaf'ni /, "مرسي" /mursi /, and "مدبولي" /mad'bu:li / are unique to Egypt while the names "أحمد" /ʔħ'mad/ "محمد" /muħam'mad/ can not be assigned to specific geographic region since they share the top very high frequency in all regions in the Arabic speaking countries among other names like "علي" /ʕli /. The module also gives some hints "gist" about gender and guesses on the religion for some non-Arabic origin names e.g. "جرجس" / girgis/, "مينا" /mi:na/, and "حنا" /ħan'na/ which denote Coptic or Christian names common in Egypt and Iraq. MAPS uses this embedded module to give high and reliable results. Please refer Personal Names Indexer for details.

Names indexser