
Winnow이란?
Winnow compresses RAG prompts before they hit your LLM, cutting token costs 50%+ while preserving meaning. Uses question-guided filtering + LLMLingua-2 for semantic accuracy. Key features: • FastAPI server with OpenAI-compatible proxy • Batch compression API • Question-aware filtering keeps answer-relevant tokens • Docker self-hosting, pip-installable SDK • MIT licensed
데모
스크린샷
?
아직 댓글이 없어요. 가장 먼저 남겨보세요!
Winnow에 대한 X의 실제 대화
X에 게시
