Am încercat să găsesc o metodă de a descoperi cuvintele uitate / neglijate ale limbii române, relativ ieșite din uz, evitând însă termenii foarte vechi care și-au pierdut de tot relevanța.
Cuvinte apărute în dicționare, dar care sunt întâlnite rar sau deloc în româna modernă. Suveranism lexical. Use it or lose it.
Sau cum a zics odată un robot: A computational linguistics tool to identify "forgotten" Romanian words - terms that exist in official dictionaries but have fallen out of modern usage.
🐵 ☠️ 🤗 slove zăuitate boace oțioase suveranism lexical
<style> .blink-text { animation:1.7s blinker linear infinite; color: blue; font-size: 1.25rem; } .blink-text.blink-2 { animation:2s blinker2 ease-in-out infinite; color: green; } .blink-text.blink-3 { animation:1.5s blinker2 ease-in-out infinite; color: yellow; text-shadow: 1px 1px 1px rgba(0,0,0,.4); } @keyframes blinker { 0% { opacity: .8; } 50% { opacity: 0.0; } 100% { opacity: .7; } } @keyframes blinker2 { 10% { opacity: .6; } 50% { opacity: 0.0; } 90% { opacity: .8; } } </style>- Ranks Romanian dictionary words that may merit rediscovery
- Compares the dexonline dictionary collection against selected corpus evidence
- Identifies linguistic "dark matter" - words that exist in dictionaries but have fallen out of active use
- Produces curated lists with rarity scores and linguistic metadata
Vezi și: initial specs
Voroave ranks Romanian dictionary words that may merit rediscovery. Its verdicts describe evidence in selected corpora, not proven extinction or speaker recognition. Dexonline is a community dictionary collection, not a single official dictionary.
Start with scripts and setup. For delegated fixes, use the audit handoff. Task status lives in BACKLOG.md; activity history records completed work. Read AGENTS.md before changes. CLAUDE.md carries the same project guidance.
flowchart TD
DEX[Dexonline MySQL dump] --> EX[Dictionary extractors]
EX --> LEX[Lexemes and curated candidates]
EX --> INF[Inflected forms]
EX --> DEF[Definitions and dictionary sources]
WS[Wikisource] --> HIST[Historical panel]
LU[LUMRO] --> HIST
CX[CulturaX] --> MOD[Modern panel]
LEX --> V[Paradigm aggregation and verdicts]
INF --> V
HIST --> V
MOD --> V
V --> SCORE[Shortlist and two seams]
SCORE --> BUILD[UI database builder]
DEF --> BUILD
DEX --> MEAN[Structured meanings]
MEAN --> BUILD
BUILD --> UI[public/data/ui.db]
UI --> APP[PHP + HTMX + vanilla JS]
APP --> USER[Private user-data database]
CoRoLa and subtitles are reference datasets, excluded from the current production panels.
Both frequency-screen scripts are standalone; neither feeds make_shortlist.py or ui.db.
The synonym writing aid (/sinonime) moved to a separate project, ~/devbox/sinonime, on 2026-10-06.
It reads this checkout's data/ as input; see its README.
Required: Python 3 (tested with 3.14), PHP 8.1+ (extensions pdo_sqlite, sqlite3, mbstring, json),
Node 22+, and the sqlite3 command-line tool.
Generated data (public/data/ui.db) and the source dump must be built separately.
python3 -m venv .venv
.venv/bin/pip install -r requirements-test.txt # test packages only
npm ci # jsdom and Playwright, pinned in package-lock.json
npx playwright install chromium # one-time browser download
python3 tools/run_tests.py # the full strict checkrequirements.txt is the full pipeline set (corpus, frequency screens, scrapers).
It includes the optional editable -e ../gov2/wrodfreq line, which needs that sibling checkout.
No test needs it. The PHP application needs none of these packages.
tools/run_tests.py is the single required command. It does the following:
- It stages a copy of
public/in a temp directory with a temp private directory and a random admin token. The runner never opensprivate/app.db,private/secret.keyorpublic/data/*.dbfor writing. It compares their size and mtime before and after the run. - It starts and stops its own PHP server with the dev router on a free port.
- It runs the Python tests and every
tests/test_*.jssuite. - It fails on a nonzero exit, a timeout, a missing prerequisite, or a SKIP line (a required skip).
- It reports portable suites (no built data) and artifact-dependent suites (built
ui.db) separately.
python3 tools/run_tests.py --portable-only runs the portable suites only. It is a partial check.
--list shows the suite table. To add a suite, add one Suite(...) line to SUITES in the runner.
Plain pytest runs current tests only (pytest.ini). Archived Flask tests have an explicit command:
.venv/bin/python -m pytest archive -q (six known failures).
Apache rewrite rules (public/.htaccess) are not tested here; the runner uses the dev router.
See F07.
The historical panel totals about 19.4M tokens: Wikisource 14.3M and LUMRO 5.1M.
The modern panel is CulturaX, about 17.0B tokens. These are the audited build's input totals.
Different sizes create different sampling uncertainty; normalized frequencies still share the same units.
The current pipeline uses paradigm-level occurrence gates, not the old shared ppm absence gate.
Modern thresholds rescale with the modern panel size through scaled_modern_thresholds().
| Verdict | Current validator condition |
|---|---|
extinct |
Historical occurrences ≥3, historical document evidence ≥2, modern occurrences =0 |
historical_only |
Historical evidence passes those floors; modern occurrences >0 and below the scaled rare threshold |
declining |
Historical evidence passes; modern occurrences reach the scaled rare threshold but remain below the alive threshold |
absent |
Historical evidence fails; modern occurrences remain below the alive threshold |
alive / emerging |
Modern occurrences reach the alive threshold; excluded from the shortlist |
Baseline rare/alive thresholds are 500/2,000 at 16,969,999,321 modern tokens.
emerging versus alive additionally uses rank shift. The four shortlisted verdicts do not use rank shift as their gate.
absent can include some corpus occurrences. It means insufficient historical footing, not “the most forgotten.”
Occurrence and document values are estimates after ambiguous forms are split.
Within a corpus, document evidence is a maximum over contributing forms; across panels it is summed.
LUMRO's unit is distinct authors; Wikisource's unit is pages.
Lexeme.frequency is treated as lexicographic prominence, not measured usage frequency.
That interpretation is empirical; do not turn it into an established definition of dexonline's field.
Zero is treated as missing data. Existing rarity-bin names are legacy metadata, not calibrated probabilities.
Wordfreq's zero means no useful estimate at this vocabulary range; it does not prove absence.
The quality score combines historical evidence, modern rarity, dictionary prominence/coverage,
current dictionary presence, and definition availability. Its hand-set weights are not independently evaluated.
relevant and curiosity are a partition. Class flags remain separate from the score.
The default view filters classes and uses editorial demotion plus a damped community-vote blend.
Curator choices are reviewable in data/editorial.tsv; community marks live in the private database.
Snapshot of local public/data/ui.db, checked 2026-10-05 (file dated 2026-08-18):
| Measurement | Count |
|---|---|
| Words | 18,270 |
| Relevant seam | 3,499 |
| Curiosity seam | 14,771 |
| Words with structured senses | 10,798 |
| Words without a flat definition | 1,126 |
| Words with nonzero Romanian wordfreq Zipf | 38 |
These are local artifact counts, not a guarantee that the deployed artifact is identical. Structured senses can provide text where the flat definition is missing. Default-view counts change with filtering; do not copy them into undated docs.
data/word_ids.tsv: append-only permanent IDs behind compact?w=shares.data/editorial.tsv: curator picks/demotions, exported from an explicitly selected account.data/sitemap_words.tsv: tracked input to generated word-page sitemap.public/data/ui.db: generated read-only dictionary artifact.private/app.db: writable annotations, lists, devices, and game state; never upload it over an existing deployment.private/secret.key: signing key; preserve it with configuration and user-data backups.
Only public/ is deployed. Keep private storage outside the web root.
Database rebuilds currently remove the last artifact first; F08 specifies atomic replacement.
A snapshot script exists; cron, off-machine backup, and restore readiness still need verification.
- Corpus expansion and measured exclusions
- Structured senses
- Conceptual roadmap: historical critique with current-status notes
- Publication assessment: dated assessment, not an implemented evaluation
- F09 evaluation brief: next bounded research deliverable
- Archived wordfreq recipe
- Archived script guide
Current priorities are reliability and independent evaluation before feature expansion. Do not treat old plans as current commands or published outcome claims.





