Instructions to use microsoft/AesCode-32B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use microsoft/AesCode-32B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="microsoft/AesCode-32B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("microsoft/AesCode-32B") model = AutoModelForMultimodalLM.from_pretrained("microsoft/AesCode-32B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use microsoft/AesCode-32B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "microsoft/AesCode-32B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/AesCode-32B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/microsoft/AesCode-32B
- SGLang
How to use microsoft/AesCode-32B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "microsoft/AesCode-32B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/AesCode-32B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "microsoft/AesCode-32B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "microsoft/AesCode-32B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use microsoft/AesCode-32B with Docker Model Runner:
docker model run hf.co/microsoft/AesCode-32B
AesCode-32B
AesCode generates information-rich visual artifacts such as slides, posters, and dashboards as HTML/CSS. The output remains structured, editable, and verifiable, but the task poses a distinct challenge: code models cannot see how layout, hierarchy, and color come together on the canvas.
Image generators offer the opposite strength. They compose visually compelling pages but often misrender text, numbers, and logical relationships. AesCode uses an image generated from the same prompt as an aesthetic reference while following the prompt for the required content.
Reference input alone does not solve the problem. Off-the-shelf vision-language models may copy hallucinated content or ignore the reference layout. AesCode separates semantic requirements from visual cues through graph-structured supervision and decoupled cross-modal rewards.
AesCode-32B starts from Qwen3-VL-32B-Instruct and is trained with cold-start SFT followed by GDPO across seven reward channels. It is the larger released checkpoint and leads the aggregate benchmark results.
- ๐ Paper: AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards
- ๐ป Code: https://github.com/microsoft/AesCode
- ๐ค Companion model:
microsoft/AesCode-8B
Intended Use
AesCode is designed to generate information-rich visual artifacts as structured HTML/CSS. It is suited for creating editable slides, posters, dashboards, and reports, as well as for research on multimodal code generation and verifiable visual design.
Results
We evaluate on 300 infographic samples. Rule averages the programmatic Text, Bound, and Chart checks. Visual averages the Content, Layout, and Style checklist dimensions. Overall is the mean of Rule and Visual. Scores are percentages averaged over three generations per prompt with no selection.
Ref. indicates whether the model receives an image-generated reference together with the prompt. Bold marks the best score in each column and underline the second best.
| Model | Ref. | Text | Bound | Chart | Rule | Content | Layout | Style | Visual | Overall |
|---|---|---|---|---|---|---|---|---|---|---|
| GPT-5.5 | No | 88.67 | 86.18 | 82.85 | 85.90 | 79.27 | 68.09 | 52.80 | 66.72 | 76.31 |
| GPT-5.5 | Yes | 89.22 | 86.85 | 81.35 | 85.80 | 82.44 | 89.93 | 57.91 | 76.76 | 81.28 |
| Claude Opus 4.8 | No | 91.40 | 85.03 | 87.69 | 88.04 | 84.98 | 64.88 | 46.72 | 65.53 | 76.78 |
| Claude Opus 4.8 | Yes | 92.45 | 84.43 | 73.75 | 83.55 | 86.42 | 89.76 | 55.51 | 77.23 | 80.39 |
| Qwen3-VL-8B | No | 58.32 | 77.23 | 43.30 | 59.61 | 38.33 | 23.97 | 13.22 | 25.17 | 42.39 |
| Qwen3-VL-8B | Yes | 64.07 | 70.99 | 42.35 | 59.14 | 52.90 | 51.33 | 29.92 | 44.72 | 51.93 |
| Qwen3-VL-32B | No | 66.00 | 80.31 | 41.36 | 62.55 | 49.81 | 34.85 | 19.34 | 34.67 | 48.61 |
| Qwen3-VL-32B | Yes | 71.19 | 79.61 | 50.77 | 67.19 | 62.04 | 64.59 | 38.43 | 55.02 | 61.10 |
| AesCode-8B | Yes | 94.06 | 88.36 | 87.79 | 90.07 | 86.41 | 87.80 | 53.21 | 75.80 | 82.94 |
| AesCode-32B | Yes | 95.34 | 97.27 | 90.37 | 94.33 | 85.76 | 90.58 | 55.99 | 77.44 | 85.89 |
AesCode-32B improves Visual over its reference-conditioned Qwen3-VL-32B backbone by 22.4 points and Overall by 24.8. It leads the comparison on Overall at 85.89 and reaches Rule 94.33, compared with 85.90 for GPT-5.5 and 88.04 for Claude Opus 4.8.
Two results stand out. Table/Chart at 90.37 leads all models: requiring tables to use HTML table structures and charts to use ECharts constrains the solution space to executable, directly verifiable implementations. Boundary at 97.27 reflects a sharp reduction in canvas overflow: a severe boundary failure recurs on 4.3% of the 300 samples, against 34.7% for GPT-5.5.
Style is the shared ceiling for every system. It awards credit only when a design needs no further visual revision before delivery, and no model clears 60.
Usage
The model follows the Qwen3-VL chat interface and requires transformers>=4.57. Supply a
detailed content prompt and, when available, a reference image for layout and style
guidance. The weights are bf16; plan for roughly 65 GB of accelerator memory for the
parameters alone.
from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "microsoft/AesCode-32B"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, dtype="auto", device_map="auto"
)
messages = [{
"role": "user",
"content": [
{"type": "image", "image": "reference.png"},
{"type": "text", "text": "<your content prompt>"},
],
}]
inputs = processor.apply_chat_template(
messages, add_generation_prompt=True, tokenize=True,
return_dict=True, return_tensors="pt",
).to(model.device)
out = model.generate(
**inputs, do_sample=True, temperature=0.8, top_p=0.95, max_new_tokens=12000
)
print(processor.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
For batched generation, serve with vLLM and allow two images per prompt:
vllm serve microsoft/AesCode-32B --tensor-parallel-size 4 \
--limit-mm-per-prompt image=2 --max-model-len 24576
Reported scores use three rollouts per prompt at temperature 0.8, top-p 0.95, up to 12,000 output tokens and a 24,576-token context.
The model emits a complete HTML document. Tables use HTML table structures and charts use ECharts specifications, keeping both directly inspectable and verifiable.
The reference image is optional. Reference-conditioned training internalizes visual planning into the policy. Measured on AesCode-8B, withholding the reference at inference costs only 1.00 Visual point, against 19.55 for the Qwen3-VL-8B-Instruct backbone and 10.04 for GPT-5.5. Quality is still highest when a reference is supplied.
The repository provides the training prompt template and a pipeline that turns a short idea into the prompt and reference pair the model expects.
Training
Cold-start SFT. 3,000 demonstrations at learning rate 1e-5.
Reinforcement learning. GDPO over 7,408 prompts for 520 steps, using verl's FSDP-vLLM hybrid engine with no critic and no separately trained reward model. AdamW at a constant 5e-6, no warmup, weight decay 0.01, 128 prompts per step with 8 rollouts each, KL and entropy coefficients both 0.001, rollout temperature 1.0 and top-p 1.0, prompt and response each capped at 8,192 tokens. AesCode-8B uses the same recipe and a shorter 400-step run.
The reward is not a single scalar. Each target is represented as a design graph describing the whole canvas, which makes every property individually attributable. From that graph come two complementary reward families: deterministic verifiers score the properties parsable from the code and its rendering, while a VLM judge scores the non-parsable ones using a sample-specific Visual Graph Rubric tied to the graph's elements and relations. These form seven channels: execution, text, boundary, tablechart, layout, whitespace, and design. Each is normalized within its rollout group before aggregation, so a dominant signal cannot drown out a weaker one.
Candidate HTML is scored by rendering it in a sandboxed Playwright browser with external requests blocked, which exports the DOM, computed styles, bounding boxes, console status and a screenshot.
Training code, the verifier stack and the Visual Graph Rubric builder are released at https://github.com/microsoft/AesCode.
Citation
@inproceedings{aescode,
title = {AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards},
booktitle = {Under review},
year = {2027}
}
License
Released under the Apache 2.0 license, following the Qwen3-VL backbone.
- Downloads last month
- 8