codelion's picture
Memory estimates: embeddings per quant, SWA slots, vision projector; plainer descriptions
9ced3b9 verified
|
Raw History Blame Contribute Delete
3.15 kB
metadata
title: Local Model Explorer
emoji: 🦙
colorFrom: blue
colorTo: yellow
sdk: docker
app_port: 7860
pinned: false
license: mit
short_description: See which GGUF quants fit your GPU, Mac or CPU
tags:
  - gguf
  - llama.cpp
  - local-llm

Local Model Explorer

Set your GPU, Mac or CPU and see which GGUF quants of a model fit, from all uploaders, how much of each goes into VRAM and RAM, and the llama.cpp or Ollama command to run it.

What it does

  • Lists the most downloaded and trending GGUF repos on Hugging Face, grouped by the model they quantize. Modified versions (abliterated, uncensored, merges) are models of their own. MLX, ONNX, AWQ, GPTQ, FP8 and EXL2/3 builds are linked from each model.
  • Reads each model's GGUF header (layer count, attention layout, per-tensor sizes) and plans placement the way llama.cpp does: layers on the GPU with -ngl, MoE experts in RAM with --n-cpu-moe, embeddings in RAM. Sliding-window, hybrid and latent-attention KV caches are sized correctly.
  • Ranks models on fit, how much model your hardware can hold, popularity, recency and community reports. No uploader is favoured.
  • Collects anonymous visit summaries, reports and pasted llama-bench results into LocalLLaMA/local-model-explorer-data.

Memory estimates only: speed comes from community benchmarks, not a formula.

Architecture

app/          FastAPI
  catalogue.py    GGUF listing + other formats, grouping by model, 6h refresh, snapshot fallback
  gguf_header.py  GGUF metadata + tensor table from a ranged request
  geometry.py     weight split (embeddings / output / layers / experts) and KV cache per layer
  detail.py       per-repo quant files, geometry cache, shared Hub call budget
  hardware.py     GPU / unified / CPU presets
  fit.py          placement: GPU, experts in RAM, partial offload, CPU; llama.cpp and Ollama commands
  recommend.py    per-model best quant, ranking
  bench.py        llama-bench markdown / JSON parser
  events.py, sink.py, stats.py, ratelimit.py   dataset
static/       page + stats page
data/         catalogue, geometry and file-listing snapshots
scripts/      build_snapshots.py, deploy.py, compact.py

Configuration

Env var Default Meaning
EXPLORER_SINK local hub, local or off
DATASET_REPO LocalLLaMA/local-model-explorer-data dataset for hub
HF_TOKEN Secret: fine-grained write token for the dataset only
HF_READ_TOKEN Secret: read token for Hub calls (3000 instead of 500 calls per 5 minutes)
EXPLORER_HUB_CALLS 400 Hub calls allowed per 5-minute window; raise to ~2500 with a read token
EXPLORER_PREWARM 400 popular models whose file listings are read in the background

Develop

uv venv .venv --python 3.12 && uv pip install --python .venv/bin/python -r requirements.txt pytest
.venv/bin/python -m pytest
EXPLORER_SINK=local .venv/bin/uvicorn app.main:app --port 7860
python scripts/build_snapshots.py --top 800      # refresh bundled snapshots (use a read token)