{"id":1036,"date":"2025-08-29T00:47:00","date_gmt":"2025-08-29T05:47:00","guid":{"rendered":"https:\/\/www.jkspeaks.com\/wordpress\/?p=1036"},"modified":"2026-07-05T11:58:59","modified_gmt":"2026-07-05T16:58:59","slug":"running-rag-semantic-search-locally","status":"publish","type":"post","link":"https:\/\/www.jkspeaks.com\/wordpress\/ai-data\/running-rag-semantic-search-locally\/","title":{"rendered":"Running RAG &amp; Semantic Search\u2026 Locally!"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">One of the challenges I\u2019ve faced with vector databases like Chroma DB is how heavy they can get even with a handful of small PDFs. I have seen them consume 180+ MB once vectorized. Great for experimentation, but not exactly lightweight or (work) laptop\u2011friendly.<br \/><br \/>Those who have been following me, know that I am a huge fan of local, self-hosted setups. That\u2019s why I\u2019m excited about what UC Berkeley Sky Computing Lab has released with LEANN. This is a local vector index for RAG, fully compatible with Claude Code, Ollama, and GPT\u2011OSS. Meaning, you can now run semantic search and retrieval\u2011augmented generation (RAG) on your laptop with several large files without a sweat. Think of it as your own personal, self-hosted NotebookLM.<br \/><\/p>\n\n\n\n<figure class=\"wp-block-image size-full\"><a href=\"https:\/\/www.jkspeaks.com\/wordpress\/wp-content\/uploads\/2025\/10\/image.png\"><img loading=\"lazy\" decoding=\"async\" width=\"800\" height=\"436\" src=\"https:\/\/www.jkspeaks.com\/wordpress\/wp-content\/uploads\/2025\/10\/image.png\" alt=\"\" class=\"wp-image-1039\" srcset=\"https:\/\/www.jkspeaks.com\/wordpress\/wp-content\/uploads\/2025\/10\/image.png 800w, https:\/\/www.jkspeaks.com\/wordpress\/wp-content\/uploads\/2025\/10\/image-300x164.png 300w, https:\/\/www.jkspeaks.com\/wordpress\/wp-content\/uploads\/2025\/10\/image-768x419.png 768w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" \/><\/a><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\"><br \/>Given the popularity of NotebookLM, I see some clear edge use cases for LEANN. Especially for Enterprises where keeping sensitive data on-premise is sacrosanct. My instinct says, we are not too far from having work laptops pre-imaged with custom, organizationally tuned LLMs and these privately hosted vector stores for internal consumption by consultants like me while on the go.<br \/><br \/>This weekend, I plan to build a custom <a href=\"https:\/\/www.linkedin.com\/company\/langflow\/\">Langflow<\/a> component for LEANN and put it through its paces. Would you like me to share my firsthand experience and results here?<br \/><br \/>For those who want to dig in directly:<br \/>\ud83d\udd17 GitHub: <a href=\"https:\/\/github.com\/StarTrail-org\/LEANN\">https:\/\/github.com\/StarTrail-org\/LEANN<\/a><br \/>\ud83d\udcc4 Research paper: <a href=\"https:\/\/arxiv.org\/abs\/2506.08276\">https:\/\/arxiv.org\/abs\/2506.08276<\/a><br \/><br \/>\ud83d\udc49 Technology Leaders: Imagine having a portable, private, and efficient way to bring knowledge management and AI assistance closer to your teams, without needing cloud dependencies. Isn&#8217;t that game\u2011changing?<br \/><br \/>On-Premise (is gone) -> Cloud (is moving) -> Edge Device (here to stay)<\/p>\n","protected":false},"excerpt":{"rendered":"<p>One of the challenges I\u2019ve faced with vector databases like Chroma DB is how heavy they can get even with a handful of small PDFs. I have seen them consume 180+ MB once vectorized. Great for experimentation, but not exactly lightweight or (work) laptop\u2011friendly. Those who have been following me, know that I am a&#8230;<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_monsterinsights_skip_tracking":false,"_monsterinsights_sitenote_active":false,"_monsterinsights_sitenote_note":"","_monsterinsights_sitenote_category":0,"_kad_post_transparent":"","_kad_post_title":"","_kad_post_layout":"","_kad_post_sidebar_id":"","_kad_post_content_style":"","_kad_post_vertical_padding":"","_kad_post_feature":"","_kad_post_feature_position":"","_kad_post_header":false,"_kad_post_footer":false,"_kad_post_classname":"","footnotes":"","_links_to":"","_links_to_target":""},"categories":[21],"tags":[123,187,194,211,78],"class_list":["post-1036","post","type-post","status-publish","format-standard","hentry","category-ai-data","tag-ai","tag-architecture","tag-rag","tag-semanticsearch","tag-strategy"],"acf":{"phase":"3","cluster":"Technical","topics":["AI Architecture","AI Engineering","LLM","Semantic Layer"]},"_links":{"self":[{"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/posts\/1036","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/comments?post=1036"}],"version-history":[{"count":3,"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/posts\/1036\/revisions"}],"predecessor-version":[{"id":1125,"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/posts\/1036\/revisions\/1125"}],"wp:attachment":[{"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/media?parent=1036"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/categories?post=1036"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.jkspeaks.com\/wordpress\/wp-json\/wp\/v2\/tags?post=1036"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}