Prepare prerequisites
You need the MiniCPM5 2B Midtrain GGUF file named MiniCPM5-2B-F16.gguf, a llama.cpp build with the llama-server binary, and at least 8 GB of memory or a supported GPU to host the model. Obtain the GGUF file so it is available at the path you will reference when starting the server.