Instructions to use AndrewThompson1233/maba-v1-600m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AndrewThompson1233/maba-v1-600m with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AndrewThompson1233/maba-v1-600m", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AndrewThompson1233/maba-v1-600m", trust_remote_code=True, device_map="auto") - RWKV
How to use AndrewThompson1233/maba-v1-600m with RWKV:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- RecurrentGemma
How to use AndrewThompson1233/maba-v1-600m with RecurrentGemma:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AndrewThompson1233/maba-v1-600m with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AndrewThompson1233/maba-v1-600m" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AndrewThompson1233/maba-v1-600m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AndrewThompson1233/maba-v1-600m
- SGLang
How to use AndrewThompson1233/maba-v1-600m with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AndrewThompson1233/maba-v1-600m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AndrewThompson1233/maba-v1-600m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AndrewThompson1233/maba-v1-600m" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AndrewThompson1233/maba-v1-600m", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AndrewThompson1233/maba-v1-600m with Docker Model Runner:
docker model run hf.co/AndrewThompson1233/maba-v1-600m
Maba-v1-600M
Sub-quadratic hybrid causal language model with linear recurrence and sliding GQA
Architecture Repository • Benchmarks • Quickstart
Maba-v1-600M is a 752M-parameter causal language model (600M active non-embedding backbone) trained on ~150 billion tokens. It uses a 3:1 hybrid design: 75% Gated DeltaNet-2 linear recurrence layers and 25% Grouped-Query Attention layers. This structure provides O(N) context scaling and cuts KV-cache memory requirements by 75% compared to standard Transformers.
Hardware implementations, CUDA kernels, and theoretical derivations are documented in the Maba Architecture Repository.
Benchmarks
Evaluations were conducted with lm-evaluation-harness on a single Tesla T4 (16 GB VRAM) in float16 precision with batch size 4. Total runtime was 2 hours 1 minute across 104,876 test items.
Note on dual-pass inference: Dual-pass execution was disabled in this release due to a parameter bug in the training script. This was caught when training had already practically finished. All benchmark numbers reported above reflect standard single-pass inference. Dual-pass evaluation will be released in an upcoming checkpoint.
Benchmark Comparison
| Benchmark | Maba-v1-600M | Qwen3-0.6B | SmolLM2-360M | Llama-3.2-1B |
|---|---|---|---|---|
| MMLU (Academic Knowledge, 0-shot) | 49.2% | 47.2% | 35.8% | 49.3% |
| GSM8K (Math Reasoning, 5-shot) | 32.5% | 43.0% | 3.2% | 26.2% |
| ARC-Challenge (Science Reasoning, 0-shot) | 38.1% | 42.3% | 35.8% | 46.2% |
| Winogrande (Context & Logic, 0-shot) | 58.6% | 59.2% | 52.5% | 61.2% |
| HellaSwag (Common Sense, 0-shot) | 50.5% | 53.8% | 54.5% | 68.8% |
All Maba evaluations were executed via lm-evaluation-harness in float16 on a single Tesla T4 GPU.
MMLU Performance by Domain
| Domain | Accuracy |
|---|---|
| Social Sciences | 55.31% |
| Applied & Professional | 53.04% |
| STEM | 45.54% |
| Humanities | 45.14% |
Top MMLU Disciplines
| Discipline | Accuracy |
|---|---|
| Marketing | 77.78% |
| International Law | 74.38% |
| US Foreign Policy | 69.00% |
| Management | 67.96% |
| High School Psychology | 67.89% |
Specifications
| Attribute | Value |
|---|---|
| Total Parameters | 752M |
| Active Backbone Parameters | 600M |
| Layers | 24 (18 Gated DeltaNet-2 + 6 GQA) |
| Hidden Size | 1024 |
| Intermediate Size (SwiGLU) | 2816 |
| Attention Heads | 16 Query / 8 Key-Value (GQA) |
| Vocabulary Size | 248,320 |
| Context Length | Up to 32,768 tokens |
| Pretraining Volume | ~150B tokens |
Quickstart
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "AndrewThompson1233/maba-v1-600m"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto",
trust_remote_code=True
)
messages = [
{"role": "system", "content": "You are Maba, an assistant built on the Maba v1 architecture."},
{"role": "user", "content": "Explain the difference between linear attention and quadratic attention."}
]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(
**inputs,
max_new_tokens=128,
do_sample=True,
temperature=0.7,
top_p=0.9
)
response = tokenizer.decode(outputs[0][inputs.input_ids.shape[1]:], skip_special_tokens=True)
print(response)
Architecture Reference
For architecture design, training dynamics, benchmarks against quadratic baselines, and implementation details, visit the Maba Architecture Repository.
- Downloads last month
- 290
Collection including AndrewThompson1233/maba-v1-600m
Evaluation results
- Accuracy on MMLUself-reported49.210
- Exact Match on GSM8Kself-reported32.520
- Normalized Accuracy on HellaSwagself-reported50.450
- Accuracy on Winograndeself-reported58.560
- Normalized Accuracy on ARC-Challengeself-reported38.050