Major breakthrough in my mini-data center operations. Over the weekend, I reworked most of the hardware. And built my own "Open Router" in-house front-end. Thanks to Qwen-27B and a fork in llama.cpp allowing KV cache streaming from system RAM, I was able to bring online a bunch of 16GB cards to handle multiple 27B streams to achieve large-scale document processing for our AI engines.
As a result, where I used to have 48 streams running at peak, I now have HUNDREDS of streams online, using the exact same hardware. This will allow us to index and clean (normalize) more science papers, more article content, more PDFs, etc., as we continue to build out our in-house knowledge base that powers BrightAnswers.ai and BrightLearn.ai both of which are free to the public.
Enjoy!