---
description: Lightbits Inferra, a KV cache orchestration engine, claims 16x session density, 100x lower latency, and 10M-token contexts on today&#039;s GPUs.
title: Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts
image: https://www.storagereview.com/wp-content/uploads/2026/09/StorageReview-lightbits-inferra-diagram.jpg
---

 

[](//storagereview.com/) 

---

[ ](https://www.facebook.com/@storagereview) [ ](https://twitter.com/storagereview) [ ](https://www.instagram.com/storagereview/) [ ](https://www.linkedin.com/company/storagereview-com) [ ](https://www.youtube.com/user/storagereview?sub%5Fconfirmation=1) [ ](mailto:info@storagereview.com) [ ](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [ ](https://www.reddit.com/r/StorageReview/) [ ](https://discord.gg/TwMHb4azdC) [ ](https://www.storagereview.com/rss.xml) [ ](https://www.tiktok.com/@storagereview) 

[](https://www.storagereview.com) 

≡ Menu 
* [Home](https://www.storagereview.com/)
* [Storage Reviews](https://www.storagereview.com/review)  
  * [Consumer Reviews](https://www.storagereview.com/consumer)
  * [Enterprise Reviews](https://www.storagereview.com/enterprise)
  * [Ubiquiti Reviews](https://www.storagereview.com/best/ubiquiti-reviews)
* [SR Merch](https://store.storagereview.com)
* [Leaderboards](https://www.storagereview.com/best)  
  * [Best Storage Arrays](https://www.storagereview.com/best/storage-arrays)
  * [Best Enterprise SSDs](https://www.storagereview.com/best/enterprise-ssds)
  * [Best Servers](https://www.storagereview.com/best/servers)
  * [Best Desktops for Local AI](https://www.storagereview.com/best/desktops-local-ai)
  * [Best Laptops for Local AI](https://www.storagereview.com/best/laptops-local-ai)
  * [Best Local LLM Tools](https://www.storagereview.com/best/local-llm-tools)
  * [Agentic AI Hardware](https://www.storagereview.com/best/agentic-ai-hardware)
  * [Best SSDs & Hard Drives](https://www.storagereview.com/best%5Fdrives)
  * [Best Portable SSDs](https://www.storagereview.com/best/portable-ssds)
  * [Best Business Laptops](https://www.storagereview.com/best/business-laptops)
  * [Best Mobile Workstations](https://www.storagereview.com/best/mobile-workstations)
  * [Best Desktop Workstations](https://www.storagereview.com/best/desktop-workstations)
  * [Best Laptop Battery Life](https://www.storagereview.com/best/laptop-battery-life)
* [Storage Reference Guide](https://www.storagereview.com/storage-reference-guide)
* [About SR](https://www.storagereview.com/about-storagereview)  
  * [StorageReview.com Sweepstakes Rules and Regulations](https://www.storagereview.com/storagereview-com-sweepstakes-rules-and-regulations)

Search 

[Home](https://www.storagereview.com/) » [News](https://www.storagereview.com/news) » Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

# Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts

by Harold Fritts on September 14, 2026 

[AI](https://www.storagereview.com/enterprise/ai) ◇ [Enterprise](https://www.storagereview.com/enterprise) 

Lightbits Labs, the company that invented NVMe over TCP, is moving into inference software with Inferra, a KV cache orchestration engine that makes its public debut tomorrow, September 15, at the AI Infra Summit in Santa Clara. The software virtualizes GPU memory across DRAM and NVMe storage tiers and turns the KV cache into a persistent data layer, so attention states for long-context and multi-session workloads don’t have to fit in HBM or be recomputed when they fall out. Lightbits claims up to 16x more concurrent inference sessions on existing GPUs, a greater than 100x latency improvement from prefetching attention states in place of recomputing them, and context windows up to 10 million tokens, and says the engine has already run in customer beta programs, including one with OVH.

![Lightbits Inferra architecture diagram with the KV cache engine sitting between GPU HBM, DRAM, and storage tiers, showing prefetch, compression, quantization, scheduling, encrypted transfer, and tenant isolation functions]() 

## Prefetch Instead of Recompute

The bottleneck Inferra targets is the one we walked through in our [KV cache offload to flash](https://www.storagereview.com/review/the-token-efficient-path-for-long-context-inference-kv-cache-offload-to-flash) piece: as context lengths grow and sessions multiply, the KV cache outgrows GPU memory, and the serving stack either evicts it and recomputes the prefill later or caps how many sessions a GPU can hold. Both paths burn GPU cycles on work that has already been done. Inferra keeps the cache alive across memory and storage tiers, and a predictive prefetcher pulls attention states back toward the GPU ahead of when the model needs them. Lightbits says that approach drives down both Time-to-First-Token (TTFT) and Time Per Output Token (TPOT), and the 100x figure it quotes is the latency improvement against recomputing those states from scratch.

![Lightbits Inferra block diagram showing its intelligent prefetcher, tiering manager, log-structured KV store, and security and isolation engines between LLM serving frameworks \(vLLM, TensorRT, SGLang\) and a standard NVMe SSD pool]() 

Lightbits’ block diagram lays out the four pieces: an intelligent prefetcher, a tiering manager, a log-structured KV store, and security and isolation engines, sitting between the serving framework above and a standard NVMe SSD pool below. The company lists vLLM, TensorRT, and SGLang as supported serving frameworks and says the engine is GPU-and SSD-agnostic. The log-structured store is what makes a flash tier practical for a cache that changes constantly, since it appends updates sequentially and keeps small random writes off the drives, and we looked at how a 3 DWPD drive holds up in that role in our [Solidigm D7-PS1030 review](https://www.storagereview.com/review/solidigm-d7-ps1030-review-3-dwpd-gen5-that-earned-its-keep-in-the-kv-cache-tier). On the multi-tenant side, Lightbits describes secure tenant isolation and intelligent cache management across shared inference infrastructure, with encrypted transfer between tiers, which is what lets a neocloud hand one KV pool to many customers with consistent SLAs.

Avigdor Willenz, co-founder and chairman of Lightbits Labs, framed the launch as a break from storage built for training pipelines. Inference “is a completely different paradigm,” he said, and “retrofitting legacy training and storage systems to solve the KV cache bottleneck simply doesn’t work. It was built for a different era. We engineered Inferra from the ground up specifically to solve the GPU efficiency problem, and it has already proven successful in customer beta programs.”

## OVH, Solidigm, and the Booth 219 Demos

OVH is the named beta customer. “With Inferra’s intelligent KV cache tiering, we were able to demonstrate substantial GPU utilization gains, paving the way to providing our customers a more scalable and cost-effective foundation for their AI agent and RAG workloads,” said Yaniv Fdida, chief product and technology officer at OVH. Solidigm ran the engine in its AI Central Lab on [D7-PS1010](https://www.storagereview.com/review/solidigm-ps1010-ssd-review) drives, and Avi Shetty, the company’s VP of ecosystem, solutions, and market enablement, said the results “showcased how network-attached storage with intelligent software layer virtualization can break through the memory wall for large context, agentic AI.” That pairing of a network-attached NVMe pool with a software tier is Lightbits’ home turf; its [LightOS block storage](https://www.storagereview.com/news/lightbits-adds-nvme-tcp-clustered-storage-solution-to-lightos) has been built on NVMe/TCP since 2019.

Ramesh Chettuvetty, senior vice president of AI product and business at Lightbits, called Inferra “the world’s first KV cache acceleration engine that eliminates idle GPUs and power context windows of up to 10M tokens on commodity hardware,” and said it “delivers instant payback and generates net positive savings from day one.” Those are vendor claims ahead of any independent testing, and the release doesn’t give pricing, packaging, or a general availability date; the engine is in customer beta today. Lightbits is running live demos of the session-density and latency results at Booth 219 during the summit, and has posted a Pod Efficiency Analyzer for teams that want to model the GPU savings against their own cluster before booking one.

### [Inferra by Lightbits](https://www.lightbitslabs.com/inferra/)

**Engage with StorageReview**

[Newsletter](https://www.storagereview.com/storage%5Fnewsletter) | [YouTube](https://www.youtube.com/user/StorageReview "Opens in a new window") | Podcast [iTunes](https://podcasts.apple.com/gb/podcast/storagereview-com-storage-reviews/id1060681115 "Opens in a new window")/[Spotify](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0 "Opens in a new window") | [Instagram](https://www.instagram.com/storagereview/ "Opens in a new window") | [Twitter](https://twitter.com/storagereview "Opens in a new window") | [TikTok](https://www.tiktok.com/@storagereview? "Opens in a new window") | [RSS Feed](https://www.storagereview.com/rss.xml)

![]()

### Harold Fritts

I have been in the tech industry since IBM created Selectric. My background, though, is writing. So I decided to get out of the pre-sales biz and return to my roots, doing a bit of writing but still being involved in technology. 

Previous post: [Second-Gen Single-Rack AWS Outposts Puts 2,688 vCPUs and 100TB of EBS in One 42U Rack](https://www.storagereview.com/news/second-gen-single-rack-aws-outposts-puts-2688-vcpus-and-100tb-of-ebs-in-one-42u-rack)

Next post: [NASA-IBM Lunar Foundation Model Goes Open Source With a 2M-Tile Dataset and 22% Lower Ice-Mapping Error](https://www.storagereview.com/news/nasa-ibm-lunar-foundation-model-goes-open-source-with-a-2m-tile-dataset-and-22-lower-ice-mapping-error)

Trusted Vendors

Products and solutions from our affiliate partners:

* [ Ubiquiti ](https://store.ui.com/us/en?a%5Faid=StorageReview "Opens in a new window")
* [ Newegg ](https://click.linksynergy.com/fs-bin/click?id=g5terNMDhz0&offerid=1207190.18&subid=0&type=4 "Opens in a new window")
* [  Amazon ](https://amzn.to/4fLJUEW "Opens in a new window")

Newsletter

Subscribe to the StorageReview newsletter to stay up to date on the latest news and reviews. We promise no spam!

1  

Leave this field empty if you’re human: 

## Advertisement

Content Categories

[ Facebook ](https://www.facebook.com/@storagereview) [ X ](https://twitter.com/storagereview) [ Instagram ](https://www.instagram.com/storagereview/) [ LinkedIn ](https://www.linkedin.com/company/storagereview-com) [ YouTube ](https://www.youtube.com/user/storagereview?sub%5Fconfirmation=1) [ Email ](mailto:info@storagereview.com) [ Spotify ](https://open.spotify.com/show/1y6VnznABhHeOSMOmbDTz0) [ Reddit ](https://www.reddit.com/r/StorageReview/) [ Discord ](https://discord.gg/TwMHb4azdC) [ RSS ](https://www.storagereview.com/rss.xml) [ TikTok ](https://www.tiktok.com/@storagereview) 

Copyright © 1998-2025 Flying Pig Ventures, LLC Cincinnati, Ohio. All rights reserved.

Manage your privacy

To provide the best experiences, we and our partners use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us and our partners to process personal data such as browsing behavior or unique IDs on this site and show (non-) personalized ads. Not consenting or withdrawing consent, may adversely affect certain features and functions.

Click below to consent to the above or make granular choices. Your choices will be applied to this site only. You can change your settings at any time, including withdrawing your consent, by using the toggles on the Cookie Policy, or by clicking on the manage consent button at the bottom of the screen.

Functional Functional Always active 

The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network. 

Preferences Preferences 

The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user. 

Statistics Statistics 

The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you. 

Marketing Marketing 

The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes. 

Statistics

Marketing

Features

Always active

Always active

* [Manage options](#)
* [Manage services](#)
* [Manage {vendor\_count} vendors](#)
* [Read more about these purposes](https://cookiedatabase.org/tcf/purposes/)

Accept Deny Manage options Save preferences [Manage options](#) 

* [{title}](#)
* [{title}](#)
* [{title}](#)

Manage your privacy

To provide the best experiences, we use technologies like cookies to store and/or access device information. Consenting to these technologies will allow us to process data such as browsing behavior or unique IDs on this site. Not consenting or withdrawing consent, may adversely affect certain features and functions.

Functional Functional Always active 

The technical storage or access is strictly necessary for the legitimate purpose of enabling the use of a specific service explicitly requested by the subscriber or user, or for the sole purpose of carrying out the transmission of a communication over an electronic communications network. 

Preferences Preferences 

The technical storage or access is necessary for the legitimate purpose of storing preferences that are not requested by the subscriber or user. 

Statistics Statistics 

The technical storage or access that is used exclusively for statistical purposes. The technical storage or access that is used exclusively for anonymous statistical purposes. Without a subpoena, voluntary compliance on the part of your Internet Service Provider, or additional records from a third party, information stored or retrieved for this purpose alone cannot usually be used to identify you. 

Marketing Marketing 

The technical storage or access is required to create user profiles to send advertising, or to track the user on a website or across several websites for similar marketing purposes. 

Statistics

Marketing

Features

Always active

Always active

* [Manage options](#)
* [Manage services](#)
* [Manage {vendor\_count} vendors](#)
* [Read more about these purposes](https://cookiedatabase.org/tcf/purposes/)

Accept Deny Manage options Save preferences [Manage options](#) 

* [{title}](#)
* [{title}](#)
* [{title}](#)

Manage consent Manage consent

```json
{"@context":"https:\/\/schema.org","@graph":[{"@type":"WebPage","@id":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts","url":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts","name":"Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts - StorageReview.com","isPartOf":{"@id":"https:\/\/www.storagereview.com\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts#primaryimage"},"image":{"@id":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts#primaryimage"},"thumbnailUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/09\/StorageReview-lightbits-inferra-diagram.jpg","datePublished":"2026-09-14T16:23:21+00:00","description":"Lightbits Inferra, a KV cache orchestration engine, claims 16x session density, 100x lower latency, and 10M-token contexts on today's GPUs.","breadcrumb":{"@id":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts#primaryimage","url":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/09\/StorageReview-lightbits-inferra-diagram.jpg","contentUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2026\/09\/StorageReview-lightbits-inferra-diagram.jpg","width":1500,"height":1037},{"@type":"BreadcrumbList","@id":"https:\/\/www.storagereview.com\/news\/lightbits-inferra-kv-cache-engine-claims-16x-session-density-and-10m-token-contexts#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.storagereview.com\/"},{"@type":"ListItem","position":2,"name":"News","item":"https:\/\/www.storagereview.com\/news"},{"@type":"ListItem","position":3,"name":"Lightbits Inferra KV Cache Engine Claims 16x Session Density and 10M-Token Contexts"}]},{"@type":"WebSite","@id":"https:\/\/www.storagereview.com\/#website","url":"https:\/\/www.storagereview.com\/","name":"StorageReview.com","description":"StorageReview.com is a leading provider of news and reviews throughout the entire IT stack - from the datacenter to the edge, and all points in between.","publisher":{"@id":"https:\/\/www.storagereview.com\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.storagereview.com\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.storagereview.com\/#organization","name":"StorageReview.com","url":"https:\/\/www.storagereview.com\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.storagereview.com\/#\/schema\/logo\/image\/","url":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2020\/02\/Storage-Reviews-2-2.png","contentUrl":"https:\/\/www.storagereview.com\/wp-content\/uploads\/2020\/02\/Storage-Reviews-2-2.png","width":344,"height":61,"caption":"StorageReview.com"},"image":{"@id":"https:\/\/www.storagereview.com\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/x.com\/storagereview","http:\/\/youtube.com\/user\/storagereview"]}]}
```
