• HN Mail
  • Subscribe

CPP

New in Llama.cpp: Decision Models
6 points | 1 comments

I got 2.2x more tokens per second from llama.cpp on Intel Arc
4 points | 0 comments

Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs
3 points | 0 comments

Llama.cpp banned him for accidentally tagging in a fork PR
1 points | 2 comments

Show HN: Reflex Engine Beats Both Llama.cpp and vLLM on Cold-Start to TTFT
1 points | 0 comments

ContextMemory – Markdown memory for your llama.cpp/vLLM server
1 points | 0 comments

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
194 points | 97 comments