Kalmasoft Databases
Overview
Kalmasoft maintains and manages the world’s largest central repository of multilingual databases. For over two decades, we have progressively collected, curated, and updated these datasets to serve as definitive reference materials.
Our mission is to power advanced linguistic support for software packages and accelerate innovation across natural language processing (NLP) and language engineering. By continuously refining our data, we ensure that developers, researchers, and enterprises always have access to the most accurate, reliable information.
Potential applications include:
- Combating the Financing of Terrorism
- Culture-sensitive CRM applications
- Compliance and Governance
- Customer Data Management
- Employment Diversity
- Electronic Health Records
- Ethnic-based marketing campaigns
- Fraud Detection
- Entity Matching
- Identity Resolution
- Immigration Control
- Intelligence Analysis
- KYC and Due Diligence
- Law Enforcement
- Name Screening
- PEP and sanctions screening
- Voters list correction
Information
All materials presented here are commercially available and can be customized and fine-tuned to meet your specific requirements.
Reference: DBASES
Total entries 250,000,000+
Last updated: 15/6/2026
Anthroponyms (personal names)
7 large datasets segmented by script. Built for SaaS, search engines, and LLM packages.
Diacritized native script with Roman transcription and gender. Frequency stats available upon request.
Optimized for NER and name scoring. Features diacritical marks and full Roman transcription.
Over 1 million records mapping 300K Arabic names into 6 European languages.
Millions of real-world names with gender and locale fields covering the entire Arab world.
Curated dataset of non-Arab names adapted into the Arabic script.
Hundreds of names used across Islamic cultures, including Turkey, India, and Persia.
3.6 million records phonetically mapping 300K original Arabic names into 12 global languages.
Massive 40-million-record database capturing all possible Roman spelling variations.
Rare, native personal names specific to individual Arabic-speaking countries.
Global name database featuring added Arabic transcriptions, gender markers, locales, and meanings.
Names with deceptive Latin spellings influenced heavily by native phonemic traits.
Multi-lingual words sharing identical spellings but differing completely in meaning and pronunciation.
Ethiopic personal names paired with semantic Arabic equivalents.
3 million real-world, Islamic-aligned names mapped across 40 countries.
Toponyms (geographical names)
Highly organized gazetteer of populated places formatted with multiple transcription systems.
Extensive global gazetteer ready for digital publishing via multiple transcription standards.
Comprehensive geographic database detailing world topography, mountains, waterways, and road networks.
Over 2 million Arabic-translated odonyms for CLIR, web scraping, and NER systems.
Named geographic features covering major oceans, continents, valleys, summits, and notable cities.
Core terminology database optimized for electronic dictionaries and machine translation (MT).
Entity Names
High-utility curated registry of famous individuals spanning over 100 countries.
Bilingual master suite covering global domains like sports, politics, and science by locale.
Native corporate and organizational signifiers built to power web crawlers and NER pipelines.
The largest electronic index of indigenous and rare regional names.
Global consumer, military, and tech brands optimized for search engines and MT.
Acronyms and Initialisms
Thousands of specialized short forms spanning aerospace, military, law, and media.
Orthographic Databases
Tagged linguistic corpus available in UTF-8, Windows 1256, or native KATS format.
Core triconsonantal root database in native script or processing-ready KATS ASCII.
Production-grade dataset of all regular conjugated verbs found in active text.
Comprehensive database of regular inflected surface nouns from real-world text.
Dictionary-scale database encompassing regular and irregular vocabulary.
Over 5,000 adapted words mapped to original English, French, and Turkish roots.
50,000+ multi-origin technical terms indexed in native script and KATS for text parsers.
Thousands of Amharic loanwords indexed across classical and Modern Standard Arabic.
Detailed linguistic mapping of thousands of Tigrinya loanwords in Arabic.
Curated database profiling hundreds of Tigre loanwords within Arabic texts.
Curated database profiling hundreds of Geez loanwords within Arabic texts.
Comprehensive dataset tracking hundreds of Persian loanwords in the Arabic language.
Comprehensive dataset tracking hundreds of Urdu loanwords in the Arabic language.
Comprehensive dataset tracking hundreds of Hindi loanwords in the Arabic language.
Thousands of Syriac loanwords tracked through classical and contemporary Arabic.
Cross-linguistic database detailing thousands of shared lexical tokens between Amharic and Syriac.
Dual-source loanword dataset optimized specifically for CLIR and advanced IR pipelines.
Thousands of shared vocabulary words bridging regional Sudanese colloquial Arabic and Ethiopic systems
Specialized dataset tracking Fulfulde, Hausa, and Wolof influences on Sudanese vocabulary.
Semantic Databases
Hundreds of native idioms mapped to clear semantic meanings and direct English parallels for MT.
Thousands of cultural proverbs paired with English equivalents for MT and Translation Memory (TMM).
Thousands of modern media and journalistic collocations with corresponding English parallels.
Linguistic index grouping distinct Arabic terms that share near-identical pronunciation and meaning.
Fauna and Flora
Under construction. Specialized terminology database for regional zoological nomenclature.
Under construction. Structured botanical dataset covering regional agricultural and plant taxonomy.
Ontology and Semantic
Under construction. Relational semantic framework built for advanced semantic web and NLP systems.
Under construction. Hierarchical verb network engineered for computational semantic processing.
Taxonomy Databases
Under construction. Categorized classification tree for automated text parsing and entity indexing.