

Reame
Self-hosted LLM inference on the hardware you already have
Reame이란?
Reame is a CPU-first LLM inference server on llama.cpp with an OpenAI-compatible API. Built for narrow, repetitive workloads on cheap hardware — a €5 VPS, a free tier, a 2-core ARM box. Its memory layer caches prompts, prefixes and past generations to disk, so request #100 costs a fraction of request #1. MIT.
스크린샷
?
아직 댓글이 없어요. 가장 먼저 남겨보세요!
Reame에 대한 X의 실제 대화
X에 게시





