• HN Mail
  • Subscribe

CPP

Faster prompt lookup drafting in llama.cpp
82 points | 12 comments

I got 2.2x more tokens per second from llama.cpp on Intel Arc
4 points | 0 comments

Llama.cpp Under the Hood
4 points | 0 comments

Transformers now runs llama.cpp quants
4 points | 0 comments

Show HN: Jeva.cpp – a llama.cpp fork with JEV-compatible API for all LLMs
3 points | 0 comments

Custom Models in Oh My Pi: vLLM, Llama.cpp, SGLang and More
2 points | 0 comments

Llama.cpp banned him for accidentally tagging in a fork PR
1 points | 2 comments