BashkirRoBERTa Token LID

Experimental ONNX word-level language identification for mixed Bashkir, Russian and Tatar messages.

Overview

This model labels each word and punctuation token in a message. It can retain language switches inside a sentence, including quoted speech, English inserts and inflected borrowed words. It includes a standalone CPU adapter, FP32 and dynamic INT8 graphs. INT8 is the default. The encoder and classification head were fine-tuned together for token-level language discrimination on conversational and code-switched text.

At a glance
Task Word-level language identification and code-switching (BA, RU, TAT, OTHER)
Default artifact onnx/model_int8.onnx
Source A Bashkir-language encoder and multilingual conversational annotations
Version / license v1.0-experimental / Apache-2.0

Contents

Files and Configurations

File Purpose Size
onnx/model_int8.onnx CPU default, dynamic INT8 82.4 MB
onnx/model_fp32.onnx FP32 reference with PyTorch parity 200.3 MB
spm_bashkir_bert_16k.model SentencePiece tokenizer 0.59 MB
runtime.py Standalone inference, Unicode offsets and windowing —
config.json Architecture, label order and runtime contract —
benchmark_summary.json Aggregate evaluation without source texts —
example.py Download and run the default model —
META.json Release passport and artifact hashes —
LICENSE Full Apache-2.0 license text —
SHA256SUMS Checksums of the public files —

Training texts, annotation batches, evaluation corpora, PyTorch checkpoints and development logs are not included.

Model Architecture

Property Value
Encoder 8 Pre-LayerNorm Transformer blocks; hidden size 640; 10 heads
Head Linear word-level language classifier
Vocabulary 16384 SentencePiece entries
Context 256 SentencePiece positions, including boundary tokens
Input / output input_ids → logits [batch, sequence, labels]
Alignment Each original token is encoded separately; prediction at its first subtoken

Examples

These outputs were generated by the bundled INT8 runtime. O tokens are omitted here for readability.

Input Word labels
Ул миңә see you tomorrow тип әйтте, значит завтра встретимся. Ул=BA, миңә=BA, see=OTHER, you=OTHER, tomorrow=OTHER, тип=BA, әйтте=BA, значит=RU, завтра=RU, встретимся=RU
Сохрани снимок, галереяға ҡарап табырбыҙ. Сохрани=RU, снимок=RU, галереяға=BA, ҡарап=BA, табырбыҙ=BA
Сестра спрашивает: «Җомга көнне киләсезме?» Сестра=RU, спрашивает=RU, Җомга=TAT, көнне=TAT, киләсезме=TAT

BA means Bashkir; RU Russian; TAT Tatar; OTHER other languages, Latin inserts, brands or code; O punctuation, standalone numbers and symbols. Borrowed stems with Bashkir grammar have target label BA under the current rubric. Shared BA/TAT words can be ambiguous.

Method

BashkirRoBERTa base encoder
  ├── 1. tokenization with word-level subtoken alignment
  ├── 2. fine-tune token classification head on mixed BA/RU/TAT corpus
  ├── 3. export FP32 ONNX inference graph
  └── 4. dynamic INT8 quantization & sliding-window inference

A pretrained Bashkir-language encoder was adapted for word-level multilingual tagging, then further fine-tuned with conversational annotations. Training includes multilingual fragments, synthetic code-switching and reviewed labels. Ambiguous words are masked during supervision; contextual variants of each new situation remain in the same training or validation partition. FP32 was exported from the selected checkpoint and dynamically quantized to INT8. The adapter preserves the original first-subword alignment. Long messages use overlapping complete-word windows and retain their tail. A single word that cannot fit in one window raises an explicit error.

Benchmark

The following figures use one complete message per CPU invocation. Strict lexical accuracy excludes neutral punctuation and standalone numbers. Exact messages require all word tokens in the message to match the reference tags.

Evaluation suite Messages Words FP32 word accuracy INT8 word accuracy INT8 exact messages
Conversational test suite 100 640 92.34% 92.50% 73 / 100
Code-switching & loanwords 40 263 92.40% 93.16% 29 / 40
Everyday dialogue & chat 40 344 93.60% 94.77% 33 / 40

Validation annotations were used for epoch selection and are reported separately in benchmark_summary.json. The suites above assess realistic multilingual conversational settings:

  • Conversational test suite: standard mixed Bashkir/Russian/Tatar conversational sentences.
  • Code-switching & loanwords: rapid language transitions, Russian borrowings with Bashkir grammar, and foreign words.
  • Everyday dialogue & chat: natural colloquial dialogue messages, informal questions, and responses.

Quality and Use

Use this release for exploratory annotation, language-span analysis and routing mixed-language messages. Send uncertain cases to review. Confidence is an uncalibrated model score. A majority vote over word tags can provide a coarse sentence label, but hides minority-language inserts and is not a separate trained classifier.

Limitations

  • Curated conversational test suites; recommended for routing, span analysis, and dataset cleaning.
  • Bashkir without its distinctive letters can be confused with Tatar; some forms are intrinsically ambiguous.
  • Russian-root words with Bashkir endings, shared Turkic words and short switches can be mislabeled.
  • High confidence can accompany errors. OTHER is a broad fallback rather than a complete language inventory.
  • Overlapping windows preserve coverage but cannot guarantee correct classification across long contexts.
  • Public runtime uses ONNX Runtime; the standard Transformers token-classification pipeline is not provided.

Related Resources

  • Bashkir LID (Binary) — fast binary language gate (ba vs non_ba) for corpus pre-filtering and OCR triage.
  • Bashkir Multiclass LID — whole-text 4-way language classifier (ba, tt, ru, other) when token-level spans are not needed.
  • BashkirRoBERTa — the base masked-language encoder for other adaptation tasks.

Usage

pip install onnxruntime sentencepiece numpy huggingface_hub
import importlib.util
from pathlib import Path
from huggingface_hub import snapshot_download

directory=Path(snapshot_download("failed09/bashkir-roberta-token-lid", allow_patterns=[
    "onnx/model_int8.onnx", "spm_bashkir_bert_16k.model", "config.json", "META.json", "runtime.py"]))
spec=importlib.util.spec_from_file_location("bashkir_token_lid_runtime",directory/"runtime.py")
runtime=importlib.util.module_from_spec(spec)
spec.loader.exec_module(runtime)
lid=runtime.BashkirTokenLID(directory)
for token in lid.predict("Ул миңә see you tomorrow тип әйтте, значит завтра встретимся."):
    print(token)

After downloading, inference runs locally. To use FP32, also download onnx/model_fp32.onnx and pass precision="fp32". Each output token includes its label, confidence and start/end offsets in the original Python Unicode string.

License

The model weights, tokenizer and runtime are distributed under Apache-2.0. The base encoder is also released under Apache-2.0. Rights in source texts remain with their respective owners; original corpora are not redistributed.

Citation

@software{failed09_bashkir_roberta_token_lid_2026,
  title = {BashkirRoBERTa Token LID},
  author = {failed09},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/failed09/bashkir-roberta-token-lid},
  note = {Experimental ONNX word-level language identification}
}

Open Bashkir Data and Sources 🐝

This release is part of an open-source effort to support the development, preservation and practical use of the Bashkir language. Other related models, datasets and tools are available on the author's Hugging Face profile.

The author does not claim ownership or authorship of the source texts or other materials used to derive this release; rights and licensing remain with the original authors, publishers and dataset providers. Source texts are not redistributed in this repository, so users should follow the licenses and attribution requirements of the relevant upstream resources.

Downloads last month
5
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support