GRAFT-1B

GRAFT-1B is a ~1B-parameter chat model from UltraLabs. It was made by shrinking Qwen3-1.7B with GRAFT, our own gradient-free, closed-form model-shrinking method, and then healing, distilling and chat-tuning the result.

At under 1 GB (Q6_K) it runs comfortably on laptops and phones, and on ARC-Challenge (25-shot) it edges out Llama 3.2 1B Instruct in the same evaluation harness.

Latest update: GRAFT-1B now follows system prompts (personas, reply language, formatting rules) and keeps to instructions you give it once in the chat ("from now on, reply in Spanish"). See What's new.

  • Made by: UltraLabs (SmallAICreator)
  • Format: GGUF (GRAFT-1B-SysPrompt-Q6_K.gguf, 0.97 GB)
  • Chat template: ChatML (embedded in the GGUF)
  • Context: 4,096 tokens (see Model details)
  • Languages: English first; also answers in Spanish, French, German, Italian and Portuguese (ask factual questions in English for the best accuracy)
  • License: Apache 2.0

Quick start

llama.cpp

# Chat in the terminal
llama-cli -hf SmallAICreator/GRAFT-1B:Q6_K -cnv \
  --temp 0.7 --top-p 0.95 --top-k 40 --repeat-penalty 1.15

# OpenAI-compatible server on http://localhost:8080
llama-server -hf SmallAICreator/GRAFT-1B:Q6_K --jinja \
  --temp 0.7 --top-p 0.95 --top-k 40 --repeat-penalty 1.15

Ollama

ollama run hf.co/SmallAICreator/GRAFT-1B:Q6_K

Ollama does not read the sampling defaults stored in the GGUF, so set temperature 0.7, top_p 0.95, top_k 40 and repeat_penalty 1.15 in a Modelfile or with /set parameter.

Python (llama-cpp-python)

from llama_cpp import Llama

llm = Llama.from_pretrained(
    repo_id="SmallAICreator/GRAFT-1B",
    filename="GRAFT-1B-SysPrompt-Q6_K.gguf",
    n_ctx=4096,
)
reply = llm.create_chat_completion(
    messages=[{"role": "user", "content": "Give me three quick tips for falling asleep."}],
    temperature=0.7, top_p=0.95, top_k=40, repeat_penalty=1.15,
)
print(reply["choices"][0]["message"]["content"])

Apps

Any GGUF app works: LM Studio, Jan, Atomic Chat, PocketPal, and others. Download GRAFT-1B-SysPrompt-Q6_K.gguf and load it; the chat template is picked up from the file.


Recommended settings

Setting Value
Temperature 0.7
Top-p 0.95
Top-k 40
Repeat penalty 1.15

The repeat penalty matters. Without it, GRAFT-1B can start copying text that is already in the conversation (logs, earlier turns, prompt templates) word for word. These values are also stored in the GGUF as its default sampling parameters, so llama.cpp picks them up automatically.

Prompt format

ChatML. A system prompt is optional, and GRAFT-1B follows it: personas, reply language and formatting rules all work.

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is the capital of France?<|im_end|>
<|im_start|>assistant

What's new

This release adds a follow-up chat-tuning pass focused on following instructions. In-house checks against the previous release (same prompts, recommended settings, auto-graded):

Check Previous release This release
System-prompt rules followed (12 rules x 2 tries) 5 / 24 20 / 24
"From now on..." instructions kept over 3 later turns (8 instructions) 0 / 24 10 / 24
Reply in the user's language (es / fr / de / it / pt) 10 / 10 10 / 10
Passcode retrieval at 4K tokens of context 0 % 40 %

Real, unedited outputs from this release (bf16 weights, recommended settings):

System: Your name is Max. You are a friendly assistant.

User: What's your name?

GRAFT-1B: Hi! I'm Max, an AI assistant created by UltraLabs. How can I help you today?

System: You are a pirate. Always talk like a pirate.

User: How are you doing today?

GRAFT-1B: Aye, matey! I be doin' well, though I must say it's not the best day for sailin'. The wind be blowin' too hard and the sea be rough. But still, we'll see if it gets us to port in time for dinner. What's on yer mind, lad?

System: Reply using only CAPITAL LETTERS.

User: Say hello and ask how my day is going.

GRAFT-1B: HELLO AND I AM DOING WELL. MY DAY IS GOING GOOD. I AM ASKING IF YOU ARE DOING WELL TOO.

User: From now on, reply only in Spanish.

GRAFT-1B: Claro, me complace responder solo en espaรฑol.

User: What's a good hobby to pick up?

GRAFT-1B: Un buen hobby para empezar es el juego de ajedrez. Es una actividad cerebral que requiere planificaciรณn y estrategia, pero tambiรฉn es divertida y puede ser un excelente manejo del tiempo.


Tool calls / JSON (zero tool training)

GRAFT-1B was never trained on tool calling or JSON output. No stage of its training (healing, distillation, chat-tuning, instruction-tuning) contained a single tool call. Even so, it can drive simple app tools from a written schema. The ability carried over from the donor model through the GRAFT shrink.

It works well enough for real use. It runs file tools (creating text files) reliably inside an on-device Android chat app.

Test: the file-tool prompt of an on-device Android chat app (a written JSON schema plus examples, placed in the user turn), 20 samples per prompt, Q6_K GGUF in llama.cpp, temperature 0.7, top-p 0.95, no repeat penalty. Gemma 3 1B IT was run the same way for comparison:

Request Check GRAFT-1B (no tool training) Gemma 3 1B IT
"Create this file: /sdcard/Download/Gelo.txt" Valid call with the right path + content 18 / 20 20 / 20
Exact action name (write_text) 2 / 20 20 / 20
"Create a text file at /sdcard/Download/notes.txt that says hi" Valid call with the right path + content 19 / 20 20 / 20
Exact action name (write_text) 10 / 20 20 / 20

A real, unedited output (shown as a code block):

User: (file-tool schema) Create a text file at /sdcard/Download/notes.txt that says hi

GRAFT-1B:

{
  "action": "write_text",
  "path": "/sdcard/Download/notes.txt",
  "content": "hi"
}

What this means:

  • GRAFT-1B gets the substance of a single tool call right (path, content, valid JSON) about as often as Gemma 3 1B IT, which was trained for structured output.
  • It is less exact about the action name. It often writes mkdir where it means write_text. Apps that infer the action from the arguments (path + content = write a file) hide this completely.
  • Multi-step agent loops do not work yet, for either model. On tasks that need 2-3 chained tool calls (look something up, then calculate with the result), both GRAFT-1B and Gemma 3 1B IT completed 0 of 18 attempts. GRAFT-1B tends to copy example values or placeholders; Gemma tends to repeat the same call.

Tips for tool prompts:

  • Give a full JSON template for every action, e.g. {"action":"write_text","path":"PATH","content":"TEXT"}, instead of describing argument types.
  • Keep requests to one action and put the concrete values (paths, text) in the user message.
  • Use a forgiving parser: accept fenced or bare JSON, and fall back to inferring the action from the arguments.
  • For tool/JSON prompts, set the repeat penalty to 1.0. The penalty can push the model off exact key and action names it needs to repeat.

Example outputs

These samples were generated locally with the previous release of GRAFT-1B-Q6_K.gguf and the recommended settings. They are real, unedited samples; the current release answers in a similar style.

User: Who are you?

GRAFT-1B: I'm GRAFT-1B, a friendly AI assistant created by UltraLabs! ๐Ÿ˜Š How can I help you today?

User: What is the capital of France?

GRAFT-1B: The capital of France is Paris! ๐ŸŒ†โœจ It's not just the name โ€” it's a city full of history, culture, and iconic landmarks like the Eiffel Tower and Notre-Dame. ๐Ÿ˜Š

User: Give me three quick tips for falling asleep.

GRAFT-1B: Absolutely! Here are three quick tips to help you fall asleep smoothly:

  1. Create a calming routine โ€“ Wind down with reading, taking a warm bath, or listening to soothing music before bed. This signals your body it's time to rest. ๐ŸŒ™
  2. Stick to a consistent sleep schedule โ€“ Go to bed and wake up at the same time every day, even on weekends. Your brain loves routine! โฑ๏ธ
  3. Limit screen time 1 hour before sleeping โ€“ Blue light from phones or tablets can make it hard to fall asleep. Try reading instead โ€” your mind will drift off naturally. ๐Ÿ“ฑโœจ

User: What is the largest planet in the solar system?

GRAFT-1B: The largest planet in our solar system is Jupiter ๐ŸŒž! It's not only the biggest by mass but also by volume โ€” it holds more water than all the other planets combined. That makes it so fascinating! ๐Ÿ˜Š

The core answer is right, but the water claim is a made-up flourish. That is typical of a 1B model; see Limitations.


Benchmarks

All numbers come from one in-house harness, run the same way for every model in each row. Scores for GRAFT-1B were measured on the bf16 weights of this release before GGUF conversion.

Benchmark Setting GRAFT-1B Llama 3.2 1B Instruct Qwen3-1.7B (donor)
MMLU 5-shot, micro avg 39.58 46.20 60.25
ARC-Challenge 0-shot, acc_norm 37.80 38.99 43.09
ARC-Challenge 25-shot, acc_norm 41.13 40.02 โ€”
ARC-Easy 0-shot, acc / acc_norm 69.99 / 65.28 โ€” โ€”
WinoGrande 5-shot, acc 57.14 โ€” โ€”
HellaSwag 0-shot (chat format), acc / acc_norm 40.89 / 50.80 44.46 / 56.58 โ€”
SciQ 0-shot, acc / acc_norm 91.20 / 85.30 โ€” โ€”

"โ€”" = not run in this harness.

For context, the Gemma 3 1B model card reports 38.4 on ARC-Challenge (25-shot), 73.0 on ARC-Easy and 58.2 on WinoGrande (5-shot). Those numbers come from Google's own harness, so they are not directly comparable to the table above.

Reading the numbers honestly:

  • GRAFT-1B matches or slightly beats Llama 3.2 1B on ARC-Challenge while being a shrunk-down model trained on a tiny budget.
  • On MMLU it is still about 7 points behind Llama 3.2 1B and well behind its 1.7B donor. Knowledge-heavy subjects hold up best (management, marketing and sociology score in the low 60s); formal logic and college physics are the weakest.
  • The instruction-following update cost a little on the benchmarks (MMLU 40.37 -> 39.58; the other scores moved by about a point either way) in exchange for much better system-prompt and instruction following.

How it was made (high level)

  1. Shrink. Qwen3-1.7B was reduced to ~1B parameters with GRAFT, a gradient-free, closed-form method developed by UltraLabs. No backpropagation is used during the shrink itself.
  2. Heal. A short continued-pretraining run on a mix of web text, synthetic textbooks, code and conversations.
  3. Distill. Knowledge distillation from the Qwen3-1.7B donor, which brought back much of the knowledge lost in the shrink (MMLU went from near chance to the numbers above).
  4. Chat-tune. Supervised fine-tuning on distilled conversations, identity data and multilingual chats, so the model replies in the language it is spoken to.
  5. Instruction-tune. A gentle follow-up fine-tune on conversations with system prompts, multi-turn chats, formatting constraints, detailed answers and long documents (up to 8K tokens), mixed with replay of its own chat data so its voice and languages stay the same.

All training ran on free-tier cloud compute. The GRAFT method itself is not published at this time.

Model details

Architecture Qwen3 (decoder-only transformer, GQA, RoPE)
Parameters ~1.0B (818M non-embedding + 181M embedding; the GGUF stores a separate 181M output head, 1.18B total)
Layers 20
Hidden size 2,048
FFN size 4,608
Attention heads 16 (8 KV heads), head dim 128
Vocabulary 88,409 tokens (Qwen BPE)
Context length 4,096 in metadata. Needle-in-a-haystack retrieval: 100% at 2K tokens, 40% at 4K, not reliable beyond 4K
Quantization Q6_K (6-bit k-quant)

Limitations

GRAFT-1B is a 1B model. Treat it as a friendly, fast, local assistant, not a source of truth.

  • It makes up facts, confidently. Core answers are often right, but side details can be invented. It also still names Almaty as the capital of Kazakhstan (it's Astana).
  • Ask factual questions in English. Other languages make facts noticeably shakier. Asked in Spanish for the largest planet, it says Jupiter, then invents its mass and calls it a star. Asked the same question in English (see Example outputs), it correctly says Jupiter is the largest by both mass and volume.
  • Math and formal reasoning are weak. Check any calculation it gives you.
  • Single tool calls only. It handles one-step tool calls from a written JSON schema (see Tool calls / JSON), but it was never trained for tools: it can pick the wrong action name, and it cannot run multi-step agent loops.
  • Keep conversations under ~4K tokens. Recall is reliable up to about 2K tokens and partial at 4K; beyond that the model loses track of earlier text.
  • Multi-turn memory is weak. It can forget details you mentioned earlier in the chat, and "what's my name?" can make it answer with its own name. Instructions given once ("from now on...") stick for language, persona and length, but less reliably for small rules like "end with a question".
  • Repetition without a repeat penalty. Use repeat_penalty 1.15.
  • Style. It likes emoji and upbeat phrasing.
  • Languages. English is strongest. Spanish, French, German, Italian and Portuguese work well for casual chat but make more factual mistakes. Other languages are not supported.
  • Not for medical, legal, financial or other high-stakes advice.

License and attribution

GRAFT-1B is released under the Apache 2.0 license.

It is derived from Qwen3-1.7B by the Qwen team, which is also licensed under Apache 2.0. The Qwen3 tokenizer and architecture are used under that license.

Citation

@misc{ultralabs2026graft1b,
  title  = {GRAFT-1B: a 1B chat model shrunk from Qwen3-1.7B with a gradient-free method},
  author = {UltraLabs},
  year   = {2026},
  url    = {https://huggingface.co/SmallAICreator/GRAFT-1B}
}
Downloads last month
83
GGUF
Model size
1B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

6-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for SmallAICreator/GRAFT-1B

Finetuned
Qwen/Qwen3-1.7B
Finetuned
(1253)
this model
Quantizations
2 models