Contrastive-LM commited on
Commit
e939398
·
verified ·
1 Parent(s): 0dea164

Keep --max-model-len 2048 so vLLM fits on a 24 GB GPU

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -48,7 +48,7 @@ bidirectional InfoNCE loss.
48
  pip install contrastive-lm
49
 
50
  # 1. encoder (Qwen3-8B embeddings)
51
- vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --runner pooling --port 8090 &
52
 
53
  # 2. API + playground at http://localhost:8700/ (fetches CLM_v0.1-8B.pt into ~/.cache/clm/)
54
  clm-serve
 
48
  pip install contrastive-lm
49
 
50
  # 1. encoder (Qwen3-8B embeddings)
51
+ vllm serve Qwen/Qwen3-8B --served-model-name qwen3-8b --runner pooling --max-model-len 2048 --port 8090 &
52
 
53
  # 2. API + playground at http://localhost:8700/ (fetches CLM_v0.1-8B.pt into ~/.cache/clm/)
54
  clm-serve