nitotm/efficient-language-detector
Fast, accurate language detection in pure PHP (mbstring required). No dependencies. Supports 60 languages and multiple database sizes/modes (array/string/bytes/disk) to balance speed vs memory, with performance comparable to C++ detectors.
What is a language detector?
It is a tool that identifies which language a text is written in. For example, detect("Gracias") returns "es" for Spanish.
Efficient language detector (Nito-ELD or ELD) is a fast and accurate natural language detection software, written 100% in PHP, with a speed comparable to fast C++ compiled detectors, and accuracy rivaling the best detectors to date.
It has no dependencies, easy installation, all it's needed is PHP with the mb extension.
ELD is also available in: Javascript (v2), C Library (v3, includes Python package & executable), and an outdated Python (v1) implementation. ELD PHP is v3.
$ composer require nitotm/efficient-language-detector
--prefer-dist will omit tests/, misc/ & benchmark/, or use --prefer-source to include everythingnitotm/efficient-language-detector:dev-main to try the last unstable changesarray, string, bytes, disksmall, medium, large, extralargeLow memory Modes
For Modes string, bytes and disk, all size databases can run with 128MB PHP default setting
disk mode with extralarge size, can run with almost no RAM as it only uses 0.5MB
string and bytes are a great choice for general use, as they are just ~2x slower than array
Fastest Mode: Array (higher memory usage)
For array Mode it is recommended to use OPcache, specially for the larger databases to reduce load times
We need to set opcache.interned_strings_buffer, opcache.memory_consumption high enough for each database
Check Databases for more info.
detect() expects a UTF-8 string and returns an object with a language property, containing an ISO 639-1 code (or other selected scheme), or 'und' for undetermined language.
use Nitotm\Eld\LanguageDetector;
$eld = new LanguageDetector(); // Default Size: 'large', Mode: 'string' (since v3.2)
$eld->detect('Hola, cómo te llamas?');
// object( language => string, scores() => array<string, float>, isReliable() => bool )
// ( language => 'es', scores() => ['es' => 0.25, 'nl' => 0.05], isReliable() => true )
$eld->detect('Hola, cómo te llamas?')->language;
// 'es'
To select database Size and Mode, or language output scheme. We can import Eld... constants to see avalible options.
use Nitotm\Eld\{LanguageDetector, EldDataFile, EldScheme, EldMode};
// LanguageDetector(databaseFile: ?string, outputFormat: ?string, mode: string)
$eld = new LanguageDetector(EldDataFile::SMALL, EldScheme::ISO639_1, EldMode::MODE_ARRAY);
// Database Size files: 'small', 'medium', 'large', 'extralarge'.
// Schemes: 'ISO639_1', 'ISO639_2T', 'ISO639_1_BCP47', 'ISO639_2T_BCP47' and 'FULL_TEXT'
// Database Modes: 'array', 'string', 'bytes', 'disk'. Check memory requirements for 'array'
// Constants are not mandatory, LanguageDetector('small'); will also work
Calling langSubset() once, will set the subset.
array mode the first call takes longer as it creates a new database, if save enabled (default), it will be loaded next time we make the same subset.string, bytes & disk, a "virtual" subset is created instantly, detect() will just remove unwanted languages before returning results.langSubset() in array mode, when creating the instance LanguageDetector(file)
string, bytes & disk, getting lower memory usage and increased speed, it is possible by manually converting an array database, using BlobDataBuilder()// It accepts any ISO codes. In "array" mode it will return a subset file name, if saved
// langSubset(languages: [], save: true, encode: true);
$eld->langSubset(['en', 'es', 'fr', 'it', 'nl', 'de']);
// Object ( success => bool, languages => ?array, error => ?string, file => ?string )
// ( success => true, languages => ['en', 'es'...], error => NULL, file => 'small_6_mfss...' )
// to remove the subset
$eld->langSubset();
// Load pre-saved subset directly, just like a default database
$eld_subset = new Nitotm\Eld\LanguageDetector('small_6_mfss5z1t', null, 'array');
// Build a binary database for modes 'string', 'bytes' & 'disk', from any 'array' database
// Memory requirements for 'array' database input apply
$eldBuilder = new Nitotm\Eld\BlobDataBuilder('large'); // or subset 'small_6_mfss5z1t'
// Create subset directly: new BlobDataBuilder('extralarge', ['en', 'es', 'de', 'it']);
$eldBuilder->buildDatabase();
// if enableTextCleanup(True), detect() removes Urls, .com domains, emails, alphanumerical...
// Not recommended, as urls & domains contain hints of a language, which might help accuracy
$eld->enableTextCleanup(true); // Default is false
// If needed, we can get info of the ELD instance: languages, database type, etc.
$eld->info();
// Change output scheme on demand
// 'ISO639_1', 'ISO639_2T', 'ISO639_1_BCP47', 'ISO639_2T_BCP47', 'FULL_TEXT'
$eld->setOutputScheme('ISO639_2T'); // returns bool true on success
There is a CLI wrapper (BETA version)
>./bin/eld --help on Linux.
>php bin/eld --help on Windows.
I compared ELD with a different variety of detectors, as there are not many in PHP.
| URL | Version | Core Language |
|---|---|---|
| https://github.com/nitotm/efficient-language-detector/ | 3.1.0 | PHP |
| https://github.com/pemistahl/lingua-py | 2.0.2 | Rust |
| https://github.com/facebookresearch/fastText | 0.9.2 | C++ |
| https://github.com/CLD2Owners/cld2 | Aug 21, 2015 | C++ |
| https://github.com/patrickschur/language-detection | 5.3.0 | PHP |
| https://github.com/wooorm/franc | 7.2.0 | Javascript |
Benchmarks:
- For Tatoeba, I limited all detectors to the 50 languages subset, making the comparison as fair as possible.
- Also, Tatoeba is not part of ELD training dataset (nor tuning), but it is for fasttext
Time execution benchmark for ELD size large ( check others sizes at more benchmarks )
bestEffort = True, as usually returns only one language, so it has a comparative disadvantage.Modes 'bytes' and 'string' are very similar, they differ on how they are load, and are just 2x slower than Array
Mode 'string' can be OPcache'd, more expensive compilation, but then instant load, 'bytes' has always a steady ~fast load
Special mention to 'disk' mode, while slower, is the fastest uncached load & detect for the larger databases
| Mode | Disk | Bytes | String | Bytes | String |
|---|---|---|---|---|---|
| Database Size option | Extralarge | Extralarge | Extralarge | Large | Large |
| File size | 39 MB | 39 MB | 39 MB | 20 MB | 20 MB |
| Memory usage | 0.4 MB | 40 MB | 40 MB | 22 MB | 22 MB |
| Memory usage Cached | 0.4 MB | 40 MB | 0.4 MB + OP | 22 MB | 0.4 MB + OP |
| Memory peak | 0.4 MB | 40 MB | 56 MB | 22 MB | 32 MB |
| Memory peak Cached | 0.4 MB | 40 MB | 0.4 MB + OP | 22 MB | 0.4 MB + OP |
| OPcache used memory | - | - | 39 MB | - | 20 MB |
| OPcache used interned | - | - | 0.4 MB | - | 0.4 MB |
| Load & detect() Uncached | 0.0012 sec | 0.04 sec | 0.25 sec | 0.02 sec | 0.11 sec |
| Load & detect() Cached | 0.0011 sec | 0.04 sec | 0.0003 sec | 0.02 sec | 0.0003 sec |
| Mode | Bytes | String | Bytes | String |
|---|---|---|---|---|
| Database Size option | Medium | Medium | Small | Small |
| File size | 6 MB | 6 MB | 2 MB | 2 MB |
| Memory usage | 8 MB | 8 MB | 2 MB | 2 MB |
| Memory usage Cached | 8 MB | 0.4 MB + OP | 2 MB | 0.4 MB + OP |
| Memory peak | 8 MB | 12 MB | 2 MB | 3 MB |
| Memory peak Cached | 8 MB | 0.4 MB + OP | 2 MB | 0.4 MB + OP |
| OPcache used memory | - | 0 MB | - | 0 MB |
| OPcache used interned | - | 6 MB | - | 2 MB |
| Load & detect() Uncached | 0.006 sec | 0.04 sec | 0.003 sec | 0.016 sec |
| Load & detect() Cached | 0.006 sec | 0.0003 sec | 0.002 sec | 0.0003 sec |
| Array Mode, Size: | Small | Medium | Large | Extralarge |
|---|---|---|---|---|
| Pros | Lowest memory | Equilibrated | Fastest | Most accurate |
| Cons | Least accurate | Slowest (but fast) | High memory | Highest memory |
| File size | 3 MB | 9 MB | 28 MB | 64 MB |
| Memory usage | 46 MB | 137 MB | 547 MB | 1143 MB |
| Memory usage Cached | 0.4 MB + OP | 0.4 MB + OP | 0.4 MB + OP | 0.4 MB + OP |
| Memory peak | 78 MB | 287 MB | 969 MB | 2047 MB |
| Memory peak Cached | 0.4 MB + OP | 0.4 MB + OP | 0.4 MB + OP | 0.4 MB + OP |
| OPcache used memory | 21 MB | 70 MB | 242 MB | 516 MB |
| OPcache used interned | 4 MB | 10 MB | 45 MB | 91 MB |
| Load & detect() Uncached | 0.13 sec | 0.5 sec | 1.4 sec | 3.2 sec |
| Load & detect() Cached | 0.0003 sec | 0.0003 sec | 0.0003 sec | 0.0003 sec |
| Settings (Recommended) | ||||
memory_limit |
>= 128 | >= 340 | >= 1060 | >= 2200 |
opcache.interned...* |
>= 8 (16) | >= 16 (32) | >= 60 (70) | >= 116 (128) |
opcache.memory |
>= 64 (128) | >= 128 (230) | >= 360 (450) | >= 750 (820) |
interned_strings_buffer as buffers overflow error might delay server response.opcache.interned_strings_buffer should be a minimum of 160MB (170MB).opcache.memory_consumption includes opcache.interned_strings_buffer.
opcache.memory accordingly if you want them to be loaded instantly.
To cache all default databases comfortably you would want to set it at 1200MB.Default composer install might not include these files. Use --prefer-source to include them.
new Nitotm\Eld\Tests\TestsAutoload();
$ php efficient-language-detector/tests/tests.php # Update path
benchmark/bench.php file.'und' for undeterminedoutputScheme: 'ISO639_1'am, ar, az, be, bg, bn, ca, cs, da, de, el, en, es, et, eu, fa, fi, fr, gu, he, hi, hr, hu, hy, is, it, ja, ka, kn, ko, ku, lo, lt, lv, ml, mr, ms, nl, no, or, pa, pl, pt, ro, ru, sk, sl, sq, sr, sv, ta, te, th, tl, tr, uk, ur, vi, yo, zh
outputScheme: 'FULL_TEXT'Amharic, Arabic, Azerbaijani (Latin), Belarusian, Bulgarian, Bengali, Catalan, Czech, Danish, German, Greek, English, Spanish, Estonian, Basque, Persian, Finnish, French, Gujarati, Hebrew, Hindi, Croatian, Hungarian, Armenian, Icelandic, Italian, Japanese, Georgian, Kannada, Korean, Kurdish (Arabic), Lao, Lithuanian, Latvian, Malayalam, Marathi, Malay (Latin), Dutch, Norwegian, Oriya, Punjabi, Polish, Portuguese, Romanian, Russian, Slovak, Slovene, Albanian, Serbian (Cyrillic), Swedish, Tamil, Telugu, Thai, Tagalog, Turkish, Ukrainian, Urdu, Vietnamese, Yoruba, Chinese
outputScheme: 'ISO639_1_BCP47'am, ar, az-Latn, be, bg, bn, ca, cs, da, de, el, en, es, et, eu, fa, fi, fr, gu, he, hi, hr, hu, hy, is, it, ja, ka, kn, ko, ku-Arab, lo, lt, lv, ml, mr, ms-Latn, nl, no, or, pa, pl, pt, ro, ru, sk, sl, sq, sr-Cyrl, sv, ta, te, th, tl, tr, uk, ur, vi, yo, zh
outputScheme: 'ISO639_2T'. Also available with BCP 47 ISO639_2T_BCP47amh, ara, aze, bel, bul, ben, cat, ces, dan, deu, ell, eng, spa, est, eus, fas, fin, fra, guj, heb, hin, hrv, hun, hye, isl, ita, jpn, kat, kan, kor, kur, lao, lit, lav, mal, mar, msa, nld, nor, ori, pan, pol, por, ron, rus, slk, slv, sqi, srp, swe, tam, tel, tha, tgl, tur, ukr, urd, vie, yor, zho
If you wish to donate for open source improvements, hire me for private modifications, request alternative dataset training, or contact me, please use the following link: https://linktr.ee/nitotm
How can I help you explore Laravel packages today?