Research & Benchmark Release NVIDIA RTX 4070 Verified

High-Throughput Asynchronous RLHF
with In-VRAM Tensor-Native Rewards

Decoupling autoregressive generation from policy gradient updates. Eliminates CPU string SerDes overhead via zero-copy GPU tensor rewards and guarantees bounded off-policy staleness with M2PO & GRPO.

8.05M
GRPO Tokens / sec
411k
Buffer Samples / sec
62.1%
Variance Reduction (M2PO)
41 / 41
Unit & E2E Tests Passing

Interactive Tensor-Native Pattern Matching

Simulates the GPU sliding-window convolution (.unfold()) executed directly on token ID tensors inside VRAM without CPU string decoding.

Token ID 99 is EOS. Tokens after EOS are automatically masked out.

Tensor Memory Representation Reward: 1.0
Token Sequence Tape (GPU VRAM):
Unfolded Sliding Window Match Details:
PCIe Data Bus Transfer: 0.0 Bytes (Zero Copy) Latency: < 0.05 ms