HN Mail
Subscribe
CPP
Faster prompt lookup drafting in llama.cpp
89 points
|
12 comments
I got 2.2x more tokens per second from llama.cpp on Intel Arc
4 points
|
0 comments
Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs
3 points
|
0 comments
Show HN: Llama.cpp fork with 2-4x multiGPU speed for MoE models bigger than VRAM
2 points
|
3 comments
Custom Models in Oh My Pi: vLLM, Llama.cpp, SGLang and More
2 points
|
0 comments
Llama.cpp banned him for accidentally tagging in a fork PR
1 points
|
2 comments
Show HN: Reflex Engine Beats Both Llama.cpp and vLLM on Cold-Start to TTFT
1 points
|
0 comments