BashkirRoBERTa Token LID
Experimental ONNX word-level language identification for mixed Bashkir, Russian and Tatar messages.
Overview
This model labels each word and punctuation token in a message. It can retain language switches inside a sentence, including quoted speech, English inserts and inflected borrowed words. It includes a standalone CPU adapter, FP32 and dynamic INT8 graphs. INT8 is the default. The encoder and classification head were fine-tuned together for token-level language discrimination on conversational and code-switched text.
| At a glance | |
|---|---|
| Task | Word-level language identification and code-switching (BA, RU, TAT, OTHER) |
| Default artifact | onnx/model_int8.onnx |
| Source | A Bashkir-language encoder and multilingual conversational annotations |
| Version / license | v1.0-experimental / Apache-2.0 |
Contents
Files and Configurations
| File | Purpose | Size |
|---|---|---|
onnx/model_int8.onnx |
CPU default, dynamic INT8 | 82.4 MB |
onnx/model_fp32.onnx |
FP32 reference with PyTorch parity | 200.3 MB |
spm_bashkir_bert_16k.model |
SentencePiece tokenizer | 0.59 MB |
runtime.py |
Standalone inference, Unicode offsets and windowing | — |
config.json |
Architecture, label order and runtime contract | — |
benchmark_summary.json |
Aggregate evaluation without source texts | — |
example.py |
Download and run the default model | — |
META.json |
Release passport and artifact hashes | — |
LICENSE |
Full Apache-2.0 license text | — |
SHA256SUMS |
Checksums of the public files | — |
Training texts, annotation batches, evaluation corpora, PyTorch checkpoints and development logs are not included.
Model Architecture
| Property | Value |
|---|---|
| Encoder | 8 Pre-LayerNorm Transformer blocks; hidden size 640; 10 heads |
| Head | Linear word-level language classifier |
| Vocabulary | 16384 SentencePiece entries |
| Context | 256 SentencePiece positions, including boundary tokens |
| Input / output | input_ids → logits [batch, sequence, labels] |
| Alignment | Each original token is encoded separately; prediction at its first subtoken |
Examples
These outputs were generated by the bundled INT8 runtime. O tokens are omitted here for readability.
| Input | Word labels |
|---|---|
| Ул миңә see you tomorrow тип әйтте, значит завтра встретимся. | Ул=BA, миңә=BA, see=OTHER, you=OTHER, tomorrow=OTHER, тип=BA, әйтте=BA, значит=RU, завтра=RU, встретимся=RU |
| Сохрани снимок, галереяға ҡарап табырбыҙ. | Сохрани=RU, снимок=RU, галереяға=BA, ҡарап=BA, табырбыҙ=BA |
| Сестра спрашивает: «Җомга көнне киләсезме?» | Сестра=RU, спрашивает=RU, Җомга=TAT, көнне=TAT, киләсезме=TAT |
BA means Bashkir; RU Russian; TAT Tatar; OTHER other languages, Latin inserts, brands or code; O punctuation, standalone numbers and symbols. Borrowed stems with Bashkir grammar have target label BA under the current rubric. Shared BA/TAT words can be ambiguous.
Method
BashkirRoBERTa base encoder
├── 1. tokenization with word-level subtoken alignment
├── 2. fine-tune token classification head on mixed BA/RU/TAT corpus
├── 3. export FP32 ONNX inference graph
└── 4. dynamic INT8 quantization & sliding-window inference
A pretrained Bashkir-language encoder was adapted for word-level multilingual tagging, then further fine-tuned with conversational annotations. Training includes multilingual fragments, synthetic code-switching and reviewed labels. Ambiguous words are masked during supervision; contextual variants of each new situation remain in the same training or validation partition. FP32 was exported from the selected checkpoint and dynamically quantized to INT8. The adapter preserves the original first-subword alignment. Long messages use overlapping complete-word windows and retain their tail. A single word that cannot fit in one window raises an explicit error.
Benchmark
The following figures use one complete message per CPU invocation. Strict lexical accuracy excludes neutral punctuation and standalone numbers. Exact messages require all word tokens in the message to match the reference tags.
| Evaluation suite | Messages | Words | FP32 word accuracy | INT8 word accuracy | INT8 exact messages |
|---|---|---|---|---|---|
| Conversational test suite | 100 | 640 | 92.34% | 92.50% | 73 / 100 |
| Code-switching & loanwords | 40 | 263 | 92.40% | 93.16% | 29 / 40 |
| Everyday dialogue & chat | 40 | 344 | 93.60% | 94.77% | 33 / 40 |
Validation annotations were used for epoch selection and are reported separately in benchmark_summary.json. The suites above assess realistic multilingual conversational settings:
- Conversational test suite: standard mixed Bashkir/Russian/Tatar conversational sentences.
- Code-switching & loanwords: rapid language transitions, Russian borrowings with Bashkir grammar, and foreign words.
- Everyday dialogue & chat: natural colloquial dialogue messages, informal questions, and responses.
Quality and Use
Use this release for exploratory annotation, language-span analysis and routing mixed-language messages. Send uncertain cases to review. Confidence is an uncalibrated model score. A majority vote over word tags can provide a coarse sentence label, but hides minority-language inserts and is not a separate trained classifier.
Limitations
- Curated conversational test suites; recommended for routing, span analysis, and dataset cleaning.
- Bashkir without its distinctive letters can be confused with Tatar; some forms are intrinsically ambiguous.
- Russian-root words with Bashkir endings, shared Turkic words and short switches can be mislabeled.
- High confidence can accompany errors. OTHER is a broad fallback rather than a complete language inventory.
- Overlapping windows preserve coverage but cannot guarantee correct classification across long contexts.
- Public runtime uses ONNX Runtime; the standard Transformers token-classification pipeline is not provided.
Related Resources
- Bashkir LID (Binary) — fast binary language gate (
bavsnon_ba) for corpus pre-filtering and OCR triage. - Bashkir Multiclass LID — whole-text 4-way language classifier (
ba,tt,ru,other) when token-level spans are not needed. - BashkirRoBERTa — the base masked-language encoder for other adaptation tasks.
Usage
pip install onnxruntime sentencepiece numpy huggingface_hub
import importlib.util
from pathlib import Path
from huggingface_hub import snapshot_download
directory=Path(snapshot_download("failed09/bashkir-roberta-token-lid", allow_patterns=[
"onnx/model_int8.onnx", "spm_bashkir_bert_16k.model", "config.json", "META.json", "runtime.py"]))
spec=importlib.util.spec_from_file_location("bashkir_token_lid_runtime",directory/"runtime.py")
runtime=importlib.util.module_from_spec(spec)
spec.loader.exec_module(runtime)
lid=runtime.BashkirTokenLID(directory)
for token in lid.predict("Ул миңә see you tomorrow тип әйтте, значит завтра встретимся."):
print(token)
After downloading, inference runs locally. To use FP32, also download onnx/model_fp32.onnx and pass precision="fp32". Each output token includes its label, confidence and start/end offsets in the original Python Unicode string.
License
The model weights, tokenizer and runtime are distributed under Apache-2.0. The base encoder is also released under Apache-2.0. Rights in source texts remain with their respective owners; original corpora are not redistributed.
Citation
@software{failed09_bashkir_roberta_token_lid_2026,
title = {BashkirRoBERTa Token LID},
author = {failed09},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/failed09/bashkir-roberta-token-lid},
note = {Experimental ONNX word-level language identification}
}
Open Bashkir Data and Sources 🐝
This release is part of an open-source effort to support the development, preservation and practical use of the Bashkir language. Other related models, datasets and tools are available on the author's Hugging Face profile.
The author does not claim ownership or authorship of the source texts or other materials used to derive this release; rights and licensing remain with the original authors, publishers and dataset providers. Source texts are not redistributed in this repository, so users should follow the licenses and attribution requirements of the relevant upstream resources.
- Downloads last month
- 5