Spaces:
Running
Running
|
Download README.md from codelion/local-model-explorer: direct link, hf CLI and curl.
- Browser
- Download file 3.15 kB
-
https://huggingface.co/spaces/codelion/local-model-explorer/resolve/main/README.md
- Command line
-
hf download hf://spaces/codelion/local-model-explorer/README.md
-
curl -L -o README.md https://huggingface.co/spaces/codelion/local-model-explorer/resolve/main/README.md
3.15 kB
metadata
title: Local Model Explorer
emoji: 🦙
colorFrom: blue
colorTo: yellow
sdk: docker
app_port: 7860
pinned: false
license: mit
short_description: See which GGUF quants fit your GPU, Mac or CPU
tags:
- gguf
- llama.cpp
- local-llm
Local Model Explorer
Set your GPU, Mac or CPU and see which GGUF quants of a model fit, from all uploaders, how much of each goes into VRAM and RAM, and the llama.cpp or Ollama command to run it.
What it does
- Lists the most downloaded and trending GGUF repos on Hugging Face, grouped by the model they quantize. Modified versions (abliterated, uncensored, merges) are models of their own. MLX, ONNX, AWQ, GPTQ, FP8 and EXL2/3 builds are linked from each model.
- Reads each model's GGUF header (layer count, attention layout, per-tensor sizes) and plans placement the way llama.cpp does: layers on the GPU with
-ngl, MoE experts in RAM with--n-cpu-moe, embeddings in RAM. Sliding-window, hybrid and latent-attention KV caches are sized correctly. - Ranks models on fit, how much model your hardware can hold, popularity, recency and community reports. No uploader is favoured.
- Collects anonymous visit summaries, reports and pasted
llama-benchresults intoLocalLLaMA/local-model-explorer-data.
Memory estimates only: speed comes from community benchmarks, not a formula.
Architecture
app/ FastAPI
catalogue.py GGUF listing + other formats, grouping by model, 6h refresh, snapshot fallback
gguf_header.py GGUF metadata + tensor table from a ranged request
geometry.py weight split (embeddings / output / layers / experts) and KV cache per layer
detail.py per-repo quant files, geometry cache, shared Hub call budget
hardware.py GPU / unified / CPU presets
fit.py placement: GPU, experts in RAM, partial offload, CPU; llama.cpp and Ollama commands
recommend.py per-model best quant, ranking
bench.py llama-bench markdown / JSON parser
events.py, sink.py, stats.py, ratelimit.py dataset
static/ page + stats page
data/ catalogue, geometry and file-listing snapshots
scripts/ build_snapshots.py, deploy.py, compact.py
Configuration
| Env var | Default | Meaning |
|---|---|---|
EXPLORER_SINK |
local |
hub, local or off |
DATASET_REPO |
LocalLLaMA/local-model-explorer-data |
dataset for hub |
HF_TOKEN |
Secret: fine-grained write token for the dataset only | |
HF_READ_TOKEN |
Secret: read token for Hub calls (3000 instead of 500 calls per 5 minutes) | |
EXPLORER_HUB_CALLS |
400 |
Hub calls allowed per 5-minute window; raise to ~2500 with a read token |
EXPLORER_PREWARM |
400 |
popular models whose file listings are read in the background |
Develop
uv venv .venv --python 3.12 && uv pip install --python .venv/bin/python -r requirements.txt pytest
.venv/bin/python -m pytest
EXPLORER_SINK=local .venv/bin/uvicorn app.main:app --port 7860
python scripts/build_snapshots.py --top 800 # refresh bundled snapshots (use a read token)