Instructions to use SmallAICreator/GRAFT-1B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use SmallAICreator/GRAFT-1B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf SmallAICreator/GRAFT-1B:Q6_K # Run inference directly in the terminal: llama cli -hf SmallAICreator/GRAFT-1B:Q6_K
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf SmallAICreator/GRAFT-1B:Q6_K # Run inference directly in the terminal: llama cli -hf SmallAICreator/GRAFT-1B:Q6_K
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf SmallAICreator/GRAFT-1B:Q6_K # Run inference directly in the terminal: ./llama-cli -hf SmallAICreator/GRAFT-1B:Q6_K
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf SmallAICreator/GRAFT-1B:Q6_K # Run inference directly in the terminal: ./build/bin/llama-cli -hf SmallAICreator/GRAFT-1B:Q6_K
Use Docker
docker model run hf.co/SmallAICreator/GRAFT-1B:Q6_K
- LM Studio
- Jan
- vLLM
How to use SmallAICreator/GRAFT-1B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SmallAICreator/GRAFT-1B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SmallAICreator/GRAFT-1B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/SmallAICreator/GRAFT-1B:Q6_K
- Ollama
How to use SmallAICreator/GRAFT-1B with Ollama:
ollama run hf.co/SmallAICreator/GRAFT-1B:Q6_K
- Unsloth Desktop
- Docker Model Runner
How to use SmallAICreator/GRAFT-1B with Docker Model Runner:
docker model run hf.co/SmallAICreator/GRAFT-1B:Q6_K
- Lemonade
How to use SmallAICreator/GRAFT-1B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull SmallAICreator/GRAFT-1B:Q6_K
Run and chat with the model
lemonade run user.GRAFT-1B-Q6_K
List all available models
lemonade list
- Atomic Chat
GRAFT-1B
GRAFT-1B is a ~1B-parameter chat model from UltraLabs. It was made by shrinking Qwen3-1.7B with GRAFT, our own gradient-free, closed-form model-shrinking method, and then healing, distilling and chat-tuning the result.
At under 1 GB (Q6_K) it runs comfortably on laptops and phones, and on ARC-Challenge (25-shot) it edges out Llama 3.2 1B Instruct in the same evaluation harness.
Latest update: GRAFT-1B now follows system prompts (personas, reply language, formatting rules) and keeps to instructions you give it once in the chat ("from now on, reply in Spanish"). See What's new.
- Made by: UltraLabs (SmallAICreator)
- Format: GGUF (
GRAFT-1B-SysPrompt-Q6_K.gguf, 0.97 GB) - Chat template: ChatML (embedded in the GGUF)
- Context: 4,096 tokens (see Model details)
- Languages: English first; also answers in Spanish, French, German, Italian and Portuguese (ask factual questions in English for the best accuracy)
- License: Apache 2.0
Quick start
llama.cpp
# Chat in the terminal
llama-cli -hf SmallAICreator/GRAFT-1B:Q6_K -cnv \
--temp 0.7 --top-p 0.95 --top-k 40 --repeat-penalty 1.15
# OpenAI-compatible server on http://localhost:8080
llama-server -hf SmallAICreator/GRAFT-1B:Q6_K --jinja \
--temp 0.7 --top-p 0.95 --top-k 40 --repeat-penalty 1.15
Ollama
ollama run hf.co/SmallAICreator/GRAFT-1B:Q6_K
Ollama does not read the sampling defaults stored in the GGUF, so set
temperature 0.7, top_p 0.95, top_k 40 and repeat_penalty 1.15 in a
Modelfile or with /set parameter.
Python (llama-cpp-python)
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="SmallAICreator/GRAFT-1B",
filename="GRAFT-1B-SysPrompt-Q6_K.gguf",
n_ctx=4096,
)
reply = llm.create_chat_completion(
messages=[{"role": "user", "content": "Give me three quick tips for falling asleep."}],
temperature=0.7, top_p=0.95, top_k=40, repeat_penalty=1.15,
)
print(reply["choices"][0]["message"]["content"])
Apps
Any GGUF app works: LM Studio, Jan, Atomic Chat, PocketPal, and others.
Download GRAFT-1B-SysPrompt-Q6_K.gguf and load it; the chat template is picked up from
the file.
Recommended settings
| Setting | Value |
|---|---|
| Temperature | 0.7 |
| Top-p | 0.95 |
| Top-k | 40 |
| Repeat penalty | 1.15 |
The repeat penalty matters. Without it, GRAFT-1B can start copying text that is already in the conversation (logs, earlier turns, prompt templates) word for word. These values are also stored in the GGUF as its default sampling parameters, so llama.cpp picks them up automatically.
Prompt format
ChatML. A system prompt is optional, and GRAFT-1B follows it: personas, reply language and formatting rules all work.
<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is the capital of France?<|im_end|>
<|im_start|>assistant
What's new
This release adds a follow-up chat-tuning pass focused on following instructions. In-house checks against the previous release (same prompts, recommended settings, auto-graded):
| Check | Previous release | This release |
|---|---|---|
| System-prompt rules followed (12 rules x 2 tries) | 5 / 24 | 20 / 24 |
| "From now on..." instructions kept over 3 later turns (8 instructions) | 0 / 24 | 10 / 24 |
| Reply in the user's language (es / fr / de / it / pt) | 10 / 10 | 10 / 10 |
| Passcode retrieval at 4K tokens of context | 0 % | 40 % |
Real, unedited outputs from this release (bf16 weights, recommended settings):
System: Your name is Max. You are a friendly assistant.
User: What's your name?
GRAFT-1B: Hi! I'm Max, an AI assistant created by UltraLabs. How can I help you today?
System: You are a pirate. Always talk like a pirate.
User: How are you doing today?
GRAFT-1B: Aye, matey! I be doin' well, though I must say it's not the best day for sailin'. The wind be blowin' too hard and the sea be rough. But still, we'll see if it gets us to port in time for dinner. What's on yer mind, lad?
System: Reply using only CAPITAL LETTERS.
User: Say hello and ask how my day is going.
GRAFT-1B: HELLO AND I AM DOING WELL. MY DAY IS GOING GOOD. I AM ASKING IF YOU ARE DOING WELL TOO.
User: From now on, reply only in Spanish.
GRAFT-1B: Claro, me complace responder solo en espaรฑol.
User: What's a good hobby to pick up?
GRAFT-1B: Un buen hobby para empezar es el juego de ajedrez. Es una actividad cerebral que requiere planificaciรณn y estrategia, pero tambiรฉn es divertida y puede ser un excelente manejo del tiempo.
Tool calls / JSON (zero tool training)
GRAFT-1B was never trained on tool calling or JSON output. No stage of its training (healing, distillation, chat-tuning, instruction-tuning) contained a single tool call. Even so, it can drive simple app tools from a written schema. The ability carried over from the donor model through the GRAFT shrink.
It works well enough for real use. It runs file tools (creating text files) reliably inside an on-device Android chat app.
Test: the file-tool prompt of an on-device Android chat app (a written JSON schema plus examples, placed in the user turn), 20 samples per prompt, Q6_K GGUF in llama.cpp, temperature 0.7, top-p 0.95, no repeat penalty. Gemma 3 1B IT was run the same way for comparison:
| Request | Check | GRAFT-1B (no tool training) | Gemma 3 1B IT |
|---|---|---|---|
| "Create this file: /sdcard/Download/Gelo.txt" | Valid call with the right path + content | 18 / 20 | 20 / 20 |
Exact action name (write_text) |
2 / 20 | 20 / 20 | |
| "Create a text file at /sdcard/Download/notes.txt that says hi" | Valid call with the right path + content | 19 / 20 | 20 / 20 |
Exact action name (write_text) |
10 / 20 | 20 / 20 |
A real, unedited output (shown as a code block):
User: (file-tool schema) Create a text file at /sdcard/Download/notes.txt that says hi
GRAFT-1B:
{ "action": "write_text", "path": "/sdcard/Download/notes.txt", "content": "hi" }
What this means:
- GRAFT-1B gets the substance of a single tool call right (path, content, valid JSON) about as often as Gemma 3 1B IT, which was trained for structured output.
- It is less exact about the action name. It often writes
mkdirwhere it meanswrite_text. Apps that infer the action from the arguments (path + content = write a file) hide this completely. - Multi-step agent loops do not work yet, for either model. On tasks that need 2-3 chained tool calls (look something up, then calculate with the result), both GRAFT-1B and Gemma 3 1B IT completed 0 of 18 attempts. GRAFT-1B tends to copy example values or placeholders; Gemma tends to repeat the same call.
Tips for tool prompts:
- Give a full JSON template for every action, e.g.
{"action":"write_text","path":"PATH","content":"TEXT"}, instead of describing argument types. - Keep requests to one action and put the concrete values (paths, text) in the user message.
- Use a forgiving parser: accept fenced or bare JSON, and fall back to inferring the action from the arguments.
- For tool/JSON prompts, set the repeat penalty to 1.0. The penalty can push the model off exact key and action names it needs to repeat.
Example outputs
These samples were generated locally with the previous release of
GRAFT-1B-Q6_K.gguf and the recommended settings. They are real, unedited
samples; the current release answers in a similar style.
User: Who are you?
GRAFT-1B: I'm GRAFT-1B, a friendly AI assistant created by UltraLabs! ๐ How can I help you today?
User: What is the capital of France?
GRAFT-1B: The capital of France is Paris! ๐โจ It's not just the name โ it's a city full of history, culture, and iconic landmarks like the Eiffel Tower and Notre-Dame. ๐
User: Give me three quick tips for falling asleep.
GRAFT-1B: Absolutely! Here are three quick tips to help you fall asleep smoothly:
- Create a calming routine โ Wind down with reading, taking a warm bath, or listening to soothing music before bed. This signals your body it's time to rest. ๐
- Stick to a consistent sleep schedule โ Go to bed and wake up at the same time every day, even on weekends. Your brain loves routine! โฑ๏ธ
- Limit screen time 1 hour before sleeping โ Blue light from phones or tablets can make it hard to fall asleep. Try reading instead โ your mind will drift off naturally. ๐ฑโจ
User: What is the largest planet in the solar system?
GRAFT-1B: The largest planet in our solar system is Jupiter ๐! It's not only the biggest by mass but also by volume โ it holds more water than all the other planets combined. That makes it so fascinating! ๐
The core answer is right, but the water claim is a made-up flourish. That is typical of a 1B model; see Limitations.
Benchmarks
All numbers come from one in-house harness, run the same way for every model in each row. Scores for GRAFT-1B were measured on the bf16 weights of this release before GGUF conversion.
| Benchmark | Setting | GRAFT-1B | Llama 3.2 1B Instruct | Qwen3-1.7B (donor) |
|---|---|---|---|---|
| MMLU | 5-shot, micro avg | 39.58 | 46.20 | 60.25 |
| ARC-Challenge | 0-shot, acc_norm | 37.80 | 38.99 | 43.09 |
| ARC-Challenge | 25-shot, acc_norm | 41.13 | 40.02 | โ |
| ARC-Easy | 0-shot, acc / acc_norm | 69.99 / 65.28 | โ | โ |
| WinoGrande | 5-shot, acc | 57.14 | โ | โ |
| HellaSwag | 0-shot (chat format), acc / acc_norm | 40.89 / 50.80 | 44.46 / 56.58 | โ |
| SciQ | 0-shot, acc / acc_norm | 91.20 / 85.30 | โ | โ |
"โ" = not run in this harness.
For context, the Gemma 3 1B model card reports 38.4 on ARC-Challenge (25-shot), 73.0 on ARC-Easy and 58.2 on WinoGrande (5-shot). Those numbers come from Google's own harness, so they are not directly comparable to the table above.
Reading the numbers honestly:
- GRAFT-1B matches or slightly beats Llama 3.2 1B on ARC-Challenge while being a shrunk-down model trained on a tiny budget.
- On MMLU it is still about 7 points behind Llama 3.2 1B and well behind its 1.7B donor. Knowledge-heavy subjects hold up best (management, marketing and sociology score in the low 60s); formal logic and college physics are the weakest.
- The instruction-following update cost a little on the benchmarks (MMLU 40.37 -> 39.58; the other scores moved by about a point either way) in exchange for much better system-prompt and instruction following.
How it was made (high level)
- Shrink. Qwen3-1.7B was reduced to ~1B parameters with GRAFT, a gradient-free, closed-form method developed by UltraLabs. No backpropagation is used during the shrink itself.
- Heal. A short continued-pretraining run on a mix of web text, synthetic textbooks, code and conversations.
- Distill. Knowledge distillation from the Qwen3-1.7B donor, which brought back much of the knowledge lost in the shrink (MMLU went from near chance to the numbers above).
- Chat-tune. Supervised fine-tuning on distilled conversations, identity data and multilingual chats, so the model replies in the language it is spoken to.
- Instruction-tune. A gentle follow-up fine-tune on conversations with system prompts, multi-turn chats, formatting constraints, detailed answers and long documents (up to 8K tokens), mixed with replay of its own chat data so its voice and languages stay the same.
All training ran on free-tier cloud compute. The GRAFT method itself is not published at this time.
Model details
| Architecture | Qwen3 (decoder-only transformer, GQA, RoPE) |
| Parameters | ~1.0B (818M non-embedding + 181M embedding; the GGUF stores a separate 181M output head, 1.18B total) |
| Layers | 20 |
| Hidden size | 2,048 |
| FFN size | 4,608 |
| Attention heads | 16 (8 KV heads), head dim 128 |
| Vocabulary | 88,409 tokens (Qwen BPE) |
| Context length | 4,096 in metadata. Needle-in-a-haystack retrieval: 100% at 2K tokens, 40% at 4K, not reliable beyond 4K |
| Quantization | Q6_K (6-bit k-quant) |
Limitations
GRAFT-1B is a 1B model. Treat it as a friendly, fast, local assistant, not a source of truth.
- It makes up facts, confidently. Core answers are often right, but side details can be invented. It also still names Almaty as the capital of Kazakhstan (it's Astana).
- Ask factual questions in English. Other languages make facts noticeably shakier. Asked in Spanish for the largest planet, it says Jupiter, then invents its mass and calls it a star. Asked the same question in English (see Example outputs), it correctly says Jupiter is the largest by both mass and volume.
- Math and formal reasoning are weak. Check any calculation it gives you.
- Single tool calls only. It handles one-step tool calls from a written JSON schema (see Tool calls / JSON), but it was never trained for tools: it can pick the wrong action name, and it cannot run multi-step agent loops.
- Keep conversations under ~4K tokens. Recall is reliable up to about 2K tokens and partial at 4K; beyond that the model loses track of earlier text.
- Multi-turn memory is weak. It can forget details you mentioned earlier in the chat, and "what's my name?" can make it answer with its own name. Instructions given once ("from now on...") stick for language, persona and length, but less reliably for small rules like "end with a question".
- Repetition without a repeat penalty. Use
repeat_penalty 1.15. - Style. It likes emoji and upbeat phrasing.
- Languages. English is strongest. Spanish, French, German, Italian and Portuguese work well for casual chat but make more factual mistakes. Other languages are not supported.
- Not for medical, legal, financial or other high-stakes advice.
License and attribution
GRAFT-1B is released under the Apache 2.0 license.
It is derived from Qwen3-1.7B by the Qwen team, which is also licensed under Apache 2.0. The Qwen3 tokenizer and architecture are used under that license.
Citation
@misc{ultralabs2026graft1b,
title = {GRAFT-1B: a 1B chat model shrunk from Qwen3-1.7B with a gradient-free method},
author = {UltraLabs},
year = {2026},
url = {https://huggingface.co/SmallAICreator/GRAFT-1B}
}
- Downloads last month
- 83
6-bit