

LLMSlim v0.3.0
Cut LLM token costs 40-70% with offline prompt compression
LLMSlim v0.3.0이란?
v0.3.0 ships hybrid prompt compression: offline TF-IDF extraction pre-prunes context, then an optional LLM rewrite pass semantically optimizes what remains. Result: 40-70% fewer tokens billed, sub-30ms CPU latency, and 100% retention of system directives, code blocks, and JSON schemas. Works with OpenAI, Anthropic, Gemini, LangChain, LlamaIndex, Ollama and more. Zero dependencies. Pure Python. pip install llmslim
스크린샷
?
아직 댓글이 없어요. 가장 먼저 남겨보세요!
LLMSlim v0.3.0에 대한 X의 실제 대화
X에 게시

