Text / TokenRhythm
NeoHorse 1 4B
NeoHorse 1 4B from TokenRhythm. Source-based hardware guidance from its published configuration.
EstimatedRepository opened Sep 5, 2026Source checked 9/23/2026Version: 56f0584b
LOCALRENTED GPUOPEN WEIGHTS
At a glance
- Parameters
- 4.21B
- Architecture
- qwen3_5_text
- Context length
- 262,144
- License
- apache-2.0
- Software
- vLLM
Will it run on your machine?
Save your machine to see a personalized rating and its reasoning.
Add my hardwareWays to run it
vLLM · Linux
Repository-specific command in publisher documentation; confirm dependencies and hardware in the source.
vllm serve "$MODEL_PATH" --served-model-name neohorse-1-4b --host 0.0.0.0 --port 8000 --max-model-len 262144 --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coderExplore its uses
SAFETENSORSTEXT-GENERATIONCODINGREASONING
Keep exploring
Text
Prism MLBonsai 2 · 27B
A compressed 27B-class reasoning model with publisher-provided GGUF packs and a dedicated llama.cpp fork for CUDA, Metal, and CPU.
MODEL SIZE27.36B
Text
DeepSeekDeepSeek V4.1 Flash
A multimodal reasoning model with a compressed key-value cache, published as open weights.
Text
QwenQwen3.8-Flash-Next
An experimental open-weight multimodal model with sparse attention and 262K native context.
MODEL SIZE125B language · 6B active