सटीकता और समर्थित फ़ॉर्मेट

“हर फ़ॉन्ट पर 100% सही” जैसा दावा करने के बजाय हम हर legacy format को अलग mapping और अलग tests के साथ जाँचते हैं।

ये numbers कैसे तैयार किए जाते हैं, इसकी पूरी process Testing methodology में दी गई है—profile identification, source/licence checks, differential references, round-trip tests, fuzzing और live browser QA सहित।

कृतिदेव 010 ↔ यूनिकोड देवनागरी

कृतिदेव इंजन खास तौर पर कृतिदेव 010 और यूनिकोड देवनागरी के लिए बनाया गया है। इसमें मात्राएँ, आधे अक्षर, संयुक्ताक्षर, रेफ, नुक्ता, अंक, विराम-चिह्न और लंबे टेक्स्ट ब्लॉक शामिल हैं।

  • 15,980 / 15,980 वास्तविक हिंदी शब्द canonical यूनिकोड → कृतिदेव → यूनिकोड round-trip में पास हुए।
  • 89,760 / 89,760 structured देवनागरी cases, जिनमें consonant-vowel forms, रेफ और conjunct patterns शामिल हैं, पास हुए।
  • 2,124 / 2,124 सार्वजनिक हिंदी UI/document strings mixed-text preservation checks में पास हुए।
  • 20,000 / 20,000 deterministic Hindi/English/symbol fuzz cases पास हुए।
  • लगभग 120,000 characters का live-browser round-trip बिना meaningful text loss के पूरा हुआ।

DevLys 010 ↔ यूनिकोड देवनागरी

DevLys 010 अलग user-facing profile है, लेकिन वही validated 010 Devanagari conversion core इस्तेमाल करता है। Independent DevLys 010 references से verify किया गया कि यह profile supported Kruti Dev 010 legacy character positions के साथ match करता है; कोई DevLys font binary bundle नहीं किया गया।

  • 8 / 8 DevLys 010 profile reference vectors legacy → Unicode checks में पास हुए।
  • Common reph, matra, conjunct, nukta और reverse-conversion samples ने DevLys 010 round trips पास किए।
  • Shared 010 engine की 15,980 / 15,980 real-word और 89,760 / 89,760 structured Devanagari validation coverage वही रहती है।
  • DevLys-labelled 120,000-character mixed-text round trip बिना truncation के पास हुआ।

Bamini ↔ यूनिकोड तमिल

Bamini के लिए अलग Tamil mapping engine इस्तेमाल होता है। यह कृतिदेव converter से निकाला गया mapping नहीं है। सामान्य Bamini layout sequences और documented historical aliases को अलग से जाँचा गया है।

  • एक सार्वजनिक Tamil dictionary corpus के 63,896 / 63,896 शब्द Unicode → Bamini → Unicode round-trip validation में पास हुए।
  • 300 / 300 canonical Bamini mapping entries ने individual round-trip checks पास किए।
  • 286 / 286 structured Tamil consonant-vowel forms ने canonical round-trip checks पास किए।
  • प्रकाशित Bamini reference text mwKk; murpaYk; सही रूप से அறமும் அரசியலும் में बदलता है और वापस वही legacy text देता है।

SunTommy ↔ यूनिकोड तमिल

SunTommy engine MIT-licensed public SunTommy mapping table पर आधारित है और Bamini से अलग profile के रूप में implement किया गया है। Current Azhagi documentation भी SunTommy को अलग non-Unicode Tamil encoding के रूप में सूचीबद्ध करती है।

  • 298 / 301 published source aliases अपने stated Unicode form में decode हुए। बाकी तीन aliases एक ही overloaded bare * token हैं; forward conversion source-order canonical ழூ interpretation इस्तेमाल करता है।
  • 286 / 286 structured Tamil consonant-vowel forms data-preserving रहे; 283 safe SunTommy legacy में representable थे और 3 ने Unicode-preservation fallback इस्तेमाल किया।
  • 89,401 / 89,401 ordered source-entry pairs recommended mode में data-preserving रहे; 27 pairs में token-merging रोकने के लिए invisible compatibility boundary लगी।
  • 20,000 / 20,000 mixed Tamil/English/symbol cases, 120,000-character round trip और 5,000 / 5,000 malformed legacy stability cases पास हुए।

Nudi ↔ यूनिकोड कन्नड़

Nudi इंजन Kannada Ganaka Parishat (KAGAPA) द्वारा प्रकाशित documented Kannada ASCII/Nudi-Baraha mapping और conversion rules पर आधारित है। Browser workflow के लिए reverse conversion और mixed-text preservation को अलग से validate किया गया है।

  • 95 / 95 प्रकाशित KAGAPA ASCII → Unicode reference cases पास हुए।
  • 9 / 9 प्रकाशित KAGAPA Unicode → ASCII reference cases पास हुए।
  • 685 / 685 canonical reverse-mapping entries recommended Unicode → Nudi → Unicode round trips में पास हुए।
  • 469,225 / 469,225 canonical mapping atoms के ordered pairs recommended-mode stress test में पास हुए।
  • 20,000 / 20,000 deterministic Kannada/English/symbol fuzz cases पास हुए।
  • 120,000 characters का deterministic mixed-text round trip बिना truncation या content loss के पूरा हुआ।

Chanakya ↔ यूनिकोड देवनागरी

Chanakya इंजन common legacy Chanakya profile को target करता है। Primary mapping basis एक ISC-declared public bidirectional implementation है; independent references का उपयोग differential checks और variant separation के लिए किया गया है।

  • 10 / 10 common public reference examples legacy → Unicode checks में पास हुए।
  • 201 / 201 validated canonical reverse-mapping atoms Unicode → Chanakya → Unicode round trips में पास हुए।
  • 4,576 / 4,576 consonant-vowel और reph consonant-vowel forms strict legacy round trips में पास हुए।
  • 75,504 / 75,504 two-consonant structured cases और 6,600 / 6,600 deterministic three-consonant cases recommended mode में data-preserving रहे।
  • 20,000 / 20,000 deterministic Hindi/English/symbol cases और 120,000-character mixed-text round trip पास हुआ।

Walkman-Chanakya 905 ↔ यूनिकोड देवनागरी

Walkman-Chanakya tool SIL-documented W-C-905 profile को target करता है। यह common Chanakya और Walkman-Chanakya 901 से अलग mapping है। Primary mapping MIT-licensed है और independent BSD implementation से byte mapping तथा reorder behavior cross-check किया गया है।

  • 15 / 15 SIL/BSD reference vectors legacy → Unicode checks में पास हुए।
  • 2,992 / 2,992 base consonant-vowel और reph forms recommended mode में data-preserving रहे।
  • 12,716 / 12,716 two-consonant structured forms recommended mode में data-preserving रहे।
  • 20,000 / 20,000 deterministic Hindi/English/symbol fuzz cases पास हुए।
  • 120,000-character mixed-text round trip बिना truncation के पास हुआ और 5,000 / 5,000 malformed legacy fuzz inputs बिना crash के complete हुए।

Shusha v1.0 ↔ यूनिकोड देवनागरी

Shusha engine exact SIL-documented Shusha v1.0 profile को target करता है। Forward JavaScript port SIL के multi-pass TECkit map को follow करता है और compiled official converter के against differential-test किया गया है। Reverse path legacy sequence तभी emit करता है जब वही verified forward engine original Unicode को exactly reproduce करे।

  • 256 / 256 possible single legacy bytes official SIL TECkit forward output से exact match हुए।
  • 65,536 / 65,536 possible two-byte sequences official TECkit forward output से exact match हुए।
  • 20,000 / 20,000 deterministic longer legacy byte strings official TECkit forward output से exact match हुए।
  • 3,264 / 3,264 structured consonant-vowel और reph forms data-preserving रहे और verified Shusha legacy sequences में representable थे।
  • 13,872 / 13,872 two-consonant structured forms data-preserving रहे और verified Shusha legacy sequences में representable थे।
  • 20,000 / 20,000 mixed Devanagari/English/symbol cases, 120,000-character round trip और 5,000 / 5,000 malformed legacy stability cases पास हुए।

Preeti v0.1 ↔ यूनिकोड देवनागरी

Preeti engine exact SIL-documented SAG-Preeti / Preeti Devanagari v0.1 profile को target करता है। Forward JavaScript port multi-pass TECkit map को follow करता है। Reverse path legacy sequence तभी emit करता है जब वही verified forward engine original Unicode को exactly reproduce करे।

  • 256 / 256 possible single legacy bytes SIL TECkit forward oracle से exact match हुए।
  • 65,536 / 65,536 possible two-byte sequences SIL TECkit forward oracle से exact match हुए।
  • 20,000 / 20,000 deterministic longer legacy byte strings SIL TECkit forward oracle से exact match हुए।
  • 3,264 / 3,264 structured consonant-vowel/reph forms data-preserving रहे; 3,036 strict Preeti-representable थे और 228 ने Unicode-preservation fallback इस्तेमाल किया।
  • 13,872 / 13,872 two-consonant forms data-preserving रहे; 13,035 strict Preeti-representable थे और 837 fallback थे।
  • 20,000 / 20,000 mixed Devanagari/English/symbol cases, 120,000-character round trip और 5,000 / 5,000 malformed legacy stability cases पास हुए।

GujaratiLS v1.0 ↔ यूनिकोड गुजराती

Gujarati tool exact SIL-documented GujaratiLS v1.0 profile को target करता है। Forward engine को official compiled TECkit converter के against differential-test किया गया है। Reverse engine केवल वही legacy sequence देता है जो वापस convert होने पर original Unicode exactly reproduce करे; नहीं तो recommended mode उस Gujarati run को Unicode में सुरक्षित रखता है।

  • 256 / 256 single legacy bytes official SIL TECkit output से exact match हुए।
  • 65,536 / 65,536 possible two-byte sequences official TECkit output से exact match हुए।
  • 20,000 / 20,000 deterministic longer legacy strings official TECkit forward output से exact match हुए।
  • 3,536 / 3,536 consonant-vowel/reph forms data-preserving रहे; इनमें 1,584 verified GujaratiLS legacy में representable थे और 1,952 ने Unicode-preservation fallback इस्तेमाल किया।
  • 15,028 / 15,028 two-consonant forms data-preserving रहे; 8,592 verified legacy में representable थे और 6,436 fallback थे।
  • 20,000 / 20,000 mixed Gujarati/English/symbol fuzz cases पास हुए; 1,387 cases में कम-से-कम एक fallback run था।
  • 120,000-character mixed-text round trip बिना truncation के पास हुआ और 5,000 / 5,000 malformed legacy fuzz inputs बिना crash के complete हुए।

ANU Telugu / AnupamaMedium v1.00 ↔ यूनिकोड तेलुगु

ANU Telugu engine SIL-documented “Telegu Anu” / AnupamaMedium v1.00 profile को target करता है। Forward JavaScript port को SIL के compiled TECkit converter के against differential-test किया गया है। Reverse path legacy sequence तभी emit करता है जब forward conversion original Unicode को exactly reproduce करे।

  • 65,536 / 65,536 possible two-byte legacy sequences official SIL TECkit forward output से exact match हुए।
  • 20,000 / 20,000 deterministic longer legacy byte strings official TECkit output से exact match हुए।
  • 980 / 980 structured base/reph forms data-preserving रहे; 964 strict legacy-representable थे और 16 ने Unicode-preservation fallback इस्तेमाल किया।
  • 17,150 / 17,150 two-consonant forms data-preserving रहे; 16,636 strict legacy-representable थे और 514 fallback थे।
  • 20,000 / 20,000 mixed Telugu/English/symbol fuzz cases, 120,000-character mixed-text round trip और 5,000 / 5,000 malformed-input stability cases पास हुए।

सटीकता कैसे जाँची जाती है?

इंजन permissively licensed mapping references, स्वतंत्र public implementations, public examples और regression suites से cross-check किए जाते हैं। कठिन ordering और ambiguity cases को permanent regression tests में रखा जाता है ताकि बाद के बदलाव पुराने सही behavior को न तोड़ें।

किन मामलों में सावधानी ज़रूरी है?

पुरानी font encodings सामान्य Latin character positions का भी इस्तेमाल करती हैं। इसलिए legacy text के बीच लिखा साधारण English text, मूल document formatting के बिना, हमेशा निश्चित रूप से पहचानना संभव नहीं है। URL और email के लिए अतिरिक्त protection है, और Unicode → legacy tools में mixed modern text के लिए recommended preservation mode दिया गया है।

कुछ पुराने Bamini tables में bare * एक से अधिक दुर्लभ Tamil forms के लिए इस्तेमाल हुआ है। केवल characters देखकर ऐसे बाहरी Bamini text को हमेशा निश्चित रूप से अलग करना संभव नहीं है। IndicConverter rare forms बनाते समय unambiguous canonical sequences इस्तेमाल करता है और Bamini input में overloaded bare * मिलने पर warning दिखाता है।

Recommended Unicode → Bamini mode में केवल उन edge cases पर invisible compatibility boundary डाली जाती है जहाँ दो सही Bamini tokens जुड़कर किसी दूसरे legacy token जैसे पढ़े जा सकते हैं। इससे IndicConverter के round trips lossless रहते हैं। Strict legacy mode में यह boundary नहीं जोड़ी जाती।

SunTommy की mapping अलग profile है। Published table में bare * ழூ, ஞூ, ஞு और ஙு के लिए overloaded है। IndicConverter external bare * को source-order ழூ के रूप में पढ़ता है; reverse conversion में unsafe rare forms को गलत token में बदलने के बजाय Unicode में सुरक्षित रखता है।

कुछ यूनिकोड टेक्स्ट के एक से अधिक canonically equivalent internal representations हो सकते हैं। इसलिए round-trip के बाद code-point ordering normalize हो सकती है, जबकि दिखाई देने वाला टेक्स्ट वही रहता है।

कृतिदेव के सभी variants एक जैसे नहीं हैं। हिंदी tool स्पष्ट रूप से कृतिदेव 010 के लिए है।

DevLys tool केवल DevLys 010 profile के लिए है। DevLys 020, 030 और दूसरे family variants को 010 के साथ byte-compatible नहीं माना जाता।

Nudi और दूसरे Kannada legacy workflows में version-specific या customized mappings हो सकते हैं। Kannada tool documented KAGAPA Nudi/Baraha-style mapping को follow करता है और हर legacy Kannada font के लिए universal compatibility का दावा नहीं करता।

Chanakya tool common profile के लिए है। Walkman-Chanakya 905 के लिए अब अलग W-C-905 converter है; Walkman-Chanakya 901 अभी supported profile नहीं है।

Recommended Unicode → Walkman-Chanakya 905 mode में rare unsupported Unicode run को गलत legacy text में बदलने के बजाय Unicode में रखा जा सकता है; इसे strict W-C-905 bytes नहीं माना जाता।

Shusha converter केवल Shusha v1.0 profile के लिए है। हर customized Shusha-family font को byte-compatible नहीं माना जाता। Reverse conversion legacy sequence तभी देता है जब verified Shusha v1.0 forward engine original Unicode को exactly reproduce करे; नहीं तो recommended mode उस Devanagari run को Unicode में सुरक्षित रखता है।

Preeti converter exact SAG-Preeti / Preeti Devanagari v0.1 profile के लिए है। SIL map बताता है कि historical mapping पूरी तरह reversible नहीं है और कुछ sequences equivalent conjunct/stack form में canonicalize हो सकते हैं। Reverse legacy sequence तभी emit होती है जब verified forward engine original Unicode exactly reproduce करे; नहीं तो recommended mode उस Devanagari run को Unicode में सुरक्षित रखता है।

Gujarati converter केवल GujaratiLS v1.0 profile के लिए है। Harikrishna, Gopika, LMG, Sulekh और दूसरे Gujarati legacy families को compatible नहीं माना जाता। Unsupported Gujarati sequence को recommended mode Unicode में सुरक्षित रख सकता है; ऐसे fallback को strict legacy coverage नहीं गिना जाता।

ANU Telugu converter केवल SIL के “Telegu Anu” / AnupamaMedium v1.00 profile के लिए है। इसे universal ANU Script Manager converter नहीं माना जाता। दूसरे ANU-family fonts या Priyanka/Anupama variants की mapping अलग हो सकती है। Unsupported Telugu run को recommended reverse mode Unicode में सुरक्षित रख सकता है।

सरकारी, कानूनी, शैक्षणिक, publishing या archival दस्तावेज़ में इस्तेमाल करने से पहले आउटपुट को मूल स्रोत से मिलाकर देखें।

तकनीकी आधार

Conversion logic documented या permissively licensed mapping resources पर आधारित है और independent public references से cross-check किया जाता है। Text conversion के लिए IndicConverter को proprietary font binaries bundle करने की ज़रूरत नहीं है।