ExaStor

High-Performance Storage for AI at Scale

Built for large-scale AI training and inference,
ExaStor delivers data to GPUs at high speed and with consistent performance,
maximizing AI workload performance and efficiency.

AI Infrastructure Performance Starts with Data

GPU performance alone cannot maximize the performance of AI training and inference infrastructure.

How fast and reliably massive datasets are delivered to GPUs determines AI workload throughput and GPU utilization.

GPU Idle Time
When storage cannot keep pace with GPU processing speeds, valuable GPU resources are left idle.
Data Growth and Bottlenecks
As training datasets and checkpoints grow, storage must scale not only in capacity, but also in throughput and metadata performance.
Complex Data Operations
Siloed environments for data ingestion, training, protection, and archiving increase data movement and operational complexity.

High-Performance AI Storage
for Accelerating AI/HPC Workloads

ExaStor is Gluesys’ high-performance AI storage solution, built on the Lustre parallel file system and a scale-out architecture

to deliver data to large-scale GPU clusters with high speed and consistent performance.

By distributing AI data and metadata across multiple storage targets and processing I/O from multiple GPU servers in parallel,

ExaStor reduces storage bottlenecks and GPU idle time.

From source data and training datasets to checkpoints and archives, ExaStor manages the entire AI data lifecycle

within a single global namespace. Starting with a single node, it scales capacity and performance together as data demands grow.

Maximized GPU Utilization
High-bandwidth data delivery reduces GPU idle time and improves processing efficiency for AI training and inference.
Parallel Data Processing
Lustre-based parallel file storage handles large-scale I/O from multiple GPU servers simultaneously.
Accelerated Metadata Operations
Distributed metadata and small-file optimization reduce bottlenecks in large-scale file environments.
Flexible Scalability
Start with a single node and scale capacity and throughput together as data demands grow.
Optimized Data Placement
Leverages storage tiers based on data characteristics and access patterns to optimize performance and cost.
Enterprise Data Protection
High availability, snapshots, remote replication, and data integrity verification protect AI data and ensure service continuity.
0 GiB/s
Seq. Read
0 kIOPS
Find
0 %
CPU Usage Reduction (GDS)
0 X
Data Reduction

Beyond Performance,
Toward an Enterprise AI Data Platform

High-performance Parallel File System

Allows parallel data access through Luster parallel file system, with high-speed file sharing and petabyte-level scalability through its parallel I/O architecture. Files are distributed and stored in objects across the cluster through global namespace.

NVIDIA® GPUDirect® Storage Support

Provides a storage interface to accelerate GPU applications, minimizing CPU and memory load. Maximizes GPU performance by providing a distributed file system IO driver module for applications running on GPGPU.

Flash-based Level 2 Cache

Can leverage SSD drives as secondary read cache, extending the main memory cache. When the main memory is full, the system can perform read request on the allocated cache drive, instead of using slower hard drives.

High-Speed Checkpoint Writes

ExaStor distributes large-scale checkpoint data across multiple storage nodes for parallel processing and leverages metadata caching to reduce checkpoint write time. Faster checkpointing enables more frequent checkpoints, reducing the amount of training that must be repeated after a failure and improving training continuity.

System Architecture

System Configuration

DC/SC Integrated Architecture

SC Scale-Out Architecture

Are you interested in our products?
Get in touch with Gluesys.