







SK hynix Unveils First HBF Standard Specifications with Sandisk, Presenting AI Memory Solutions at ‘FMS 2026’ | SK hynix Newsroom
▪ First HBF standard showcased within six months of consortium launch, expanding the ecosystem with participation from Google, Tenstorrent ▪ Keynote by SK hynix Executive Vice President Kim Chun-sung and Vice President Kang Uk-song on opening day, offering solutions for

2TB SanDisk Extreme PRO with USB4 | Sandisk
The SanDisk Extreme PRO with USB4 portable SSD is engineered for on-the-move professionals and modern adventurers, with turbocharged read speeds up to 3800MB/s and capacities up to 4TB.

The Memory Crisis Explained
Ash Hart on Twitter / X
MCDMA | Metal CUDA Direct Memory Access 🚀If you have a Spark and an Apple Silicon Mac, MCDMA gives you a direct RDMA path between CUDA memory and Metal-side unified memory over USB-C.Registered memory, rkeys, one-sided READ/WRITE, two-sided SEND/RECV with credit flow… pic.twitter.com/gFt8tx9xII— Ash Hart (@ashxhart) August 18, 2026

SanDisk Extreme PRO with USB4 Review: Fast, Rugged, Huge
SanDisk Extreme PRO USB4 SSD delivers top-tier speed and durability, but its large size may not suit ultra-portable needs.

The HDF5® Library & File Format - The HDF Group - ensuring long-term access and usability of HDF data and supporting users of HDF technologies
HDF5® Technical Details License: BSD-style Official Media Type: application/vnd.hdfgroup.hdf5 Standard File Extension: .h5, .hdf5 High-performance data management and storage suite Utilize the HDF5 high performance data software library and file format to manage, process, and store your heterogeneous data. HDF5 is built for fast I/O processing and storage. Download HDF5 Documentation What is HDF5®? HETEROGENEOUS DATA HDF® sup ports […]

LLM in a flash: Efficient Large Language Model Inference with Limited Memory
Large language models (LLMs) are central to modern natural language processing, delivering exceptional performance in various tasks. However, their substantial computational and memory requirements present challenges, especially for devices with limited DRAM capacity. This paper tackles the challenge of efficiently running LLMs that exceed the available DRAM capacity by storing the model parameters in flash memory, but bringing them on demand to DRAM. Our method involves constructing an inference cost model that takes into account the characteristics of flash memory, guiding us to optimize in two critical areas: reducing the volume of data transferred from flash and reading data in larger, more contiguous chunks. Within this hardware-informed framework, we introduce two principal techniques. First, "windowing" strategically reduces data transfer by reusing previously activated neurons, and second, "row-column bundling", tailored to the sequential data access strengths of flash memory, increases the size of data chunks read from flash memory. These methods collectively enable running models up to twice the size of the available DRAM, with a 4-5x and 20-25x increase in inference speed compared to naive loading approaches in CPU and GPU, respectively. Our integration of sparsity awareness, context-adaptive loading, and a hardware-oriented design paves the way for effective inference of LLMs on devices with limited memory.

balenaEtcher - Flash OS images to SD cards & USB drives
A cross-platform tool to flash OS images onto SD cards and USB drives safely and easily. Free and open source for makers around the world.

Benchmarks | EXO
Transparent benchmarks for LLMs tested on real hardware. Coming soon.
The Memory Walled Garden
The gap between first and third party memory systems

“Just” a hard drive
Dynamo | Proceedings of twenty-first ACM SIGOPS symposium on Operating systems principles
Reliability at massive scale is one of the biggest challenges we face at Amazon.com, one of the largest e-commerce operations in the world; even the slightest outage has significant financial consequences and impacts customer trust. The Amazon.com ...

Megakernel: Matching Apple Silicon Efficiency at 2x the Throughput on a RTX 3090
The first megakernel for hybrid DeltaNet/Attention LLMs. 413 tok/s at 1.87 tok/J on a 2020 RTX 3090, matching M5 Max efficiency at 1.8x throughput.

Been tinkering with what cloud functions could look like on AT Proto by storing WASM blobs. Influenced by the IPVM design of my big brain former co-workers from Fission github.com/avivash/at-functions
GitHub - avivash/at-functions
github.comBeen tinkering with what cloud functions could look like on AT Proto by storing WASM blobs. Influenced by the IPVM design of my big brain former co-workers from Fission github.com/avivash/at-functions
GitHub - avivash/at-functions
github.com