Eigen Hacks

v1 · embedding eigenbasis · axis 58

perform · faster · speed · fast · scale · benchmark

perform, faster, speed, fast, scale, benchmark, memori, high, optim, intel, nvidia, algorithm, block, stream, gpu, acceler, deep, top, power, slow, run, stop, record, big

dominant in 111630 stories · variance scale 0.0022

Energy by year

2006 → 2026

The canon

The purest instances over the years: one story per episode, by coordinate on this axis.

DateStorycoordenergy
2024-08-18CPU Performance Bottlenecks Limit Parallel Processing Speedups0.390.7
2021-10-22Arm's Next-Gen GPU Architecture Almost 200% Faster0.360.7
2013-03-29A new era of GPU benchmarking0.350.7
2011-09-15Compression Benchmarking: Size vs. Speed (I want both)0.350.7
2012-03-08Many times faster (de)compression using multiple processors.0.340.7
2024-02-25Multi-threaded computing across multiple processors demoed – promises big gains0.340.7
2024-03-17Multi-threading across multiple processors demo promises big performance gains0.340.7
2026-07-20Every Microsecond Matters:Achieving Near SpeedOfLight Latency in GPU Collectives0.340.7
2024-11-03Data movement bottlenecks to large-scale model training: Scaling past 1e28 FLOP0.340.7
2026-07-14CTA-Pipelining: A Latency-Oriented Spatial Scaling Method for Multi-GPU Systems0.340.7
2013-01-07Performance series: is thread contention bad?0.340.7
2013-07-12Battle Brewing Over Intel-ARM Benchmarking0.330.7
2020-12-02Performance benchmarking: beware frequency scaling0.330.7
2021-09-21Scaling TensorFlow to 300M predictions per second0.330.7
2025-10-27Severe performance penalty found in VSCode rendering loop0.330.7
2018-01-22Benchmarking Tensorflow Performance on Next Generation GPUs0.330.7
2025-10-13How parallelizing your builds can slow them down0.330.7
2023-07-14Nvidia Did It: Ray Tracing 10k Times Faster [video]0.330.7
2024-02-08Unleashing the Power of Preemptive Priority-Based Scheduling for RT GPU Tasks0.330.7
2011-10-30The LMAX Architecture - 100K TPS at Less than 1ms Latency0.330.7
2025-01-13Accurate benchmarking: how to account for the loop overhead0.320.7
2025-01-16Accurate benchmarking: how to account for the loop overhead0.320.7
2021-03-12Nvidia GPUs Are Much More Sensitive to CPU Performance Than AMD's0.320.7
2021-03-17Nvidia GPUs Are Much More Sensitive to CPU Performance Than AMD's0.320.7
2018-06-18Nvidia creates super slow-motion video smoother than a 300K fps camera0.320.7
2026-01-30Benchmarking with Vulkan: the curse of variable GPU clock rates0.320.7
2017-10-26Graphcore Preliminary “IPU” Benchmarks – Claims 2x to ~100x Over Nvidia V1000.320.7
2011-04-27Effect of memory bandwidth on GPU performance0.320.7
2017-04-20Benchmarking Tensorflow Performance and Cost Across Different GPU Options0.320.7
2017-05-08Benchmarking Tensorflow Performance and Cost Across Different GPU Options0.320.7
2024-05-21GhOST: a GPU Out-of-Order Scheduling Technique for Stall Reduction [pdf]0.320.7
2026-04-29Compute Scaling Will Slow Down Due to Increasing Lead Times0.320.7
2022-01-12When more Parallelism does not equal more Performance0.320.7
2023-03-13Multithreading in DCS increased FPS by 30-50%0.320.7
2014-10-01CPUs Outperform GPUs in Financial Markets Benchmark0.320.7
2016-04-10Nvidia’s Pascal GP100 GPU: massive bandwidth, enormous performance0.320.7
2019-06-11Performance Speed Limits0.320.7
2021-09-24Performance Speed Limits0.320.7
2021-09-27Performance Speed Limits0.320.7
2021-10-07Performance Speed Limits (2019)0.320.7