A Million Ticks At 60fps: Building GPU-Accelerated Trading Charts With WebGPU, LTTB Compute Shaders And Instanced Draws
Browser charting has been stuck at the same ceiling for a decade: Canvas2D dies somewhere past fifty thousand points, SVG dies far earlier, and every trading front end quietly downsamples on the CPU and hopes nobody zooms in. WebGPU broke that ceiling, and 2026's proof point is ChartGPU - an open-source library that pans and zooms a million points at 60fps, pushes 35 million points at roughly 72fps in its benchmark, and streams five million live candles while new ticks keep arriving. This is a working engineer's guide to why it works: LTTB downsampling as a compute shader, GPU hit-testing, one instanced draw call per series, the f32 timestamp precision trap that will ruin your intraday axis, and the ring-buffer upload pattern that keeps a live tick feed from stalling the render loop.
AlchmAI Engineering15 min read
1M @ 60fps
Points panned and zoomed smoothly by the open-source ChartGPU library, which drew wide attention on Hacker News in January 2026
35M points
Pushed at roughly 72fps in ChartGPU's benchmark demo - two to three orders of magnitude beyond a Canvas2D ceiling
5M candles
Chewed through by the live-streaming candlestick example while new ticks keep rolling in
1 draw call
Per series, via instanced rendering - the architectural change that actually buys the performance
We build charting systems for investment banks and trading firms, including engines that plot trillions of data points with nanosecond-level market data, so we have spent an unreasonable amount of our professional lives on this exact problem. The uncomfortable truth about browser charting until very recently is that almost every trading front end you have used is lying to you slightly. It is not drawing your data. It is drawing a CPU-downsampled approximation of your data, chosen by a heuristic written years ago, and the moment you zoom into a volatile window the approximation and the reality diverge in ways that matter to somebody's fill.
WebGPU changed the economics of this, and 2026 produced the proof point that made the rest of the industry pay attention: ChartGPU, an open-source WebGPU charting library that pans and zooms a million points at 60fps, with a benchmark demo pushing roughly 35 million points at around 72fps and a live-streaming candlestick example handling five million candles while ticks continue to arrive. It went up on Hacker News in January and the discussion was the usual mixture of delight and scepticism, which is healthy, because the interesting part is not the headline number. The interesting part is which three architectural decisions produce it, because those decisions transfer to whatever you build, whether you adopt the library or write your own engine.
Why Canvas2D And SVG Hit A Wall
It is worth being precise about the failure mode, because it determines what you have to fix. SVG dies first and dies obviously: every point is a DOM node, and the browser's layout and style machinery was never intended to hold a hundred thousand of anything. Canvas2D survives much longer because it has no retained scene graph, but it fails for a subtler reason - every lineTo and every stroke is a JavaScript call that crosses into the rasteriser, and at a million points you are paying that crossing a million times per frame on a single thread that is also running your React tree, your WebSocket handler and your tooltip logic.
The standard mitigation is CPU downsampling, usually Largest-Triangle-Three-Buckets, which is a genuinely good algorithm: it preserves visual shape - spikes, troughs, the outline a trader is actually reading - far better than naive decimation or averaging. The problem is not LTTB. The problem is where you run it. On the CPU, for a million points, you are doing a million-element pass in JavaScript on every zoom and every pan, and on a live feed on every tick batch. That pass is your frame budget, and everything else - the render, the axis, the crosshair - has to fit in what is left. Which is nothing.
LTTB As A Compute Shader
Moving LTTB to a compute shader is the single largest change, and it is conceptually simple once you see that the algorithm is embarrassingly parallel across buckets. Each output point depends on its own bucket, the previously selected point and the average of the next bucket. The classic sequential formulation threads the previous selection through the loop; the parallel formulation approximates it with the previous bucket's average, which is visually indistinguishable at the bucket counts a screen actually needs - you have at most a few thousand horizontal pixels, so a few thousand buckets - and removes the dependency chain entirely.
// One invocation per output bucket. Input is the full series resident in a
// storage buffer; output is at most one point per horizontal pixel.
struct Params {
point_count : u32,
bucket_count : u32,
range_start : u32, // index window for the current zoom
range_end : u32,
};
@group(0) @binding(0) var<storage, read> points : array<vec2<f32>>;
@group(0) @binding(1) var<storage, read_write> out : array<vec2<f32>>;
@group(0) @binding(2) var<uniform> params : Params;
fn bucket_bounds(i : u32, span : u32) -> vec2<u32> {
let size = f32(span) / f32(params.bucket_count);
let lo = params.range_start + u32(floor(f32(i) * size));
let hi = params.range_start + u32(floor(f32(i + 1u) * size));
return vec2<u32>(lo, max(hi, lo + 1u));
}
@compute @workgroup_size(64)
fn main(@builtin(global_invocation_id) gid : vec3<u32>) {
let i = gid.x;
if (i >= params.bucket_count) { return; }
let span = params.range_end - params.range_start;
let b = bucket_bounds(i, span);
// Anchor A: previous bucket's average (parallel approximation of the
// sequential "previously selected point" - visually equivalent at
// screen-resolution bucket counts, and removes the dependency chain).
var a = points[b.x];
if (i > 0u) {
let pb = bucket_bounds(i - 1u, span);
var acc = vec2<f32>(0.0, 0.0);
for (var k = pb.x; k < pb.y; k = k + 1u) { acc = acc + points[k]; }
a = acc / f32(pb.y - pb.x);
}
// Anchor C: next bucket's average.
var c = points[b.y - 1u];
if (i + 1u < params.bucket_count) {
let nb = bucket_bounds(i + 1u, span);
var acc = vec2<f32>(0.0, 0.0);
for (var k = nb.x; k < nb.y; k = k + 1u) { acc = acc + points[k]; }
c = acc / f32(nb.y - nb.x);
}
// Select the point in this bucket forming the largest triangle with A and C.
var best_area = -1.0;
var best = points[b.x];
for (var k = b.x; k < b.y; k = k + 1u) {
let p = points[k];
let area = abs((a.x - c.x) * (p.y - a.y) - (a.x - p.x) * (c.y - a.y));
if (area > best_area) { best_area = area; best = p; }
}
out[i] = best;
}One Draw Call Per Series: Instanced Candlesticks
The second big win is instancing. A candlestick is a body and a wick - conceptually two quads. Drawn naively that is two draw calls per candle, and at five thousand visible candles you have ten thousand draw calls per frame, which no amount of GPU horsepower rescues because the cost is command submission, not rasterisation. Instanced rendering inverts this: you upload one unit quad as vertex data, one instance record per candle as a storage buffer, and issue a single draw with an instance count. The vertex shader positions each instance from its record.
struct Candle {
t_offset : f32, // seconds from the window epoch - never absolute time
open : f32,
high : f32,
low : f32,
close : f32,
};
struct View {
t_min : f32, t_max : f32,
p_min : f32, p_max : f32,
half_width : f32, // half the candle body width, in normalised device units
};
@group(0) @binding(0) var<storage, read> candles : array<Candle>;
@group(0) @binding(1) var<uniform> view : View;
struct VsOut {
@builtin(position) pos : vec4<f32>,
@location(0) color : vec4<f32>,
};
// unit_quad is (0,0)..(1,1); part 0 = body, part 1 = wick, supplied as two
// instanced sub-ranges so the whole series is still a single draw call.
@vertex
fn vs(
@location(0) unit_quad : vec2<f32>,
@builtin(instance_index) inst : u32,
) -> VsOut {
let c = candles[inst / 2u];
let wick = (inst % 2u) == 1u;
let x = (c.t_offset - view.t_min) / (view.t_max - view.t_min) * 2.0 - 1.0;
let hw = select(view.half_width, view.half_width * 0.12, wick);
let top = select(max(c.open, c.close), c.high, wick);
let bot = select(min(c.open, c.close), c.low, wick);
let y_top = (top - view.p_min) / (view.p_max - view.p_min) * 2.0 - 1.0;
let y_bot = (bot - view.p_min) / (view.p_max - view.p_min) * 2.0 - 1.0;
var o : VsOut;
o.pos = vec4<f32>(
x + (unit_quad.x * 2.0 - 1.0) * hw,
mix(y_bot, y_top, unit_quad.y),
0.0, 1.0,
);
o.color = select(
vec4<f32>(0.91, 0.30, 0.36, 1.0), // down
vec4<f32>(0.13, 0.77, 0.55, 1.0), // up
c.close >= c.open,
);
return o;
}
@fragment
fn fs(in : VsOut) -> @location(0) vec4<f32> { return in.color; }One draw, two instances per candle, no per-candle JavaScript. Adding a second series is one more draw call, not a doubling of your frame time. This is the property that makes multi-instrument overlays - the thing every serious desk actually wants - go from painful to trivial.
Streaming Ticks Without Stalling The Render Loop
Static benchmarks are the easy half. A trading chart is a live object: ticks arrive continuously, the last candle mutates in place, and periodically a new candle is appended. The naive implementation reallocates and re-uploads the whole buffer on every update, which is exactly as bad as it sounds. The pattern that works is a GPU-resident ring buffer with partial writes - you allocate for the maximum window once, write only the bytes that changed, and let the vertex shader handle the wraparound.
const CANDLE_FLOATS = 5; // t_offset, o, h, l, c
const CANDLE_BYTES = CANDLE_FLOATS * 4;
export class CandleRing {
private readonly gpu: GPUBuffer;
private readonly staging = new Float32Array(CANDLE_FLOATS);
private head = 0; // next write slot
private count = 0;
constructor(private device: GPUDevice, private capacity: number) {
this.gpu = device.createBuffer({
size: capacity * CANDLE_BYTES,
usage: GPUBufferUsage.STORAGE | GPUBufferUsage.COPY_DST,
});
}
/** Mutate the in-progress candle. Called on every tick - must be cheap. */
updateLast(c: Candle): void {
if (this.count === 0) return this.append(c);
const slot = (this.head - 1 + this.capacity) % this.capacity;
this.write(slot, c);
}
/** Close the bar and start a new one. Called once per bar interval. */
append(c: Candle): void {
this.write(this.head, c);
this.head = (this.head + 1) % this.capacity;
this.count = Math.min(this.count + 1, this.capacity);
}
private write(slot: number, c: Candle): void {
// 20 bytes copied. Not a reallocation, not a full re-upload, and
// writeBuffer does not block the render loop - the copy is queued.
this.staging[0] = c.tOffset;
this.staging[1] = c.open;
this.staging[2] = c.high;
this.staging[3] = c.low;
this.staging[4] = c.close;
this.device.queue.writeBuffer(
this.gpu, slot * CANDLE_BYTES, this.staging.buffer, 0, CANDLE_BYTES,
);
}
}GPU Hit-Testing: The Stall Nobody Profiles
The third change is the least obvious and, in our experience shipping these to desks, the one users feel most. When the mouse moves over a chart, something has to answer 'which point is under the cursor?'. Done on the CPU over a million points that is a linear scan per mousemove event, and mousemove fires far more often than people assume. Even with a spatial index you are paying JavaScript on the main thread on every pointer movement, which is why so many otherwise-fast charts feel sticky under the crosshair.
The GPU answer is to render a picking pass - the same geometry, with the instance index encoded as colour, into an offscreen single-pixel-region texture around the cursor - and read back one pixel. Read-back has latency, which sounds fatal until you notice the crosshair is allowed to be one frame behind. It is a tooltip, not an order confirmation. One frame of lag is imperceptible; a 40ms main-thread scan is not.
Where This Fits Next To TradingView
A fair question at this point: why build any of this when TradingView exists? Our honest answer, having integrated the standard libraries and written fully custom engines for the same clients, is that these are complementary rather than competing choices, and the decision is about what the chart is for.
- Use the industry-standard library when the value is in the ecosystem: hundreds of indicators, drawing tools traders already know, saved layouts, symbol search. Rebuilding that is years of work with no differentiation at the end of it, and traders genuinely do not want your version of a Fibonacci retracement tool.
- Build custom when the value is in the data volume or the domain: full-depth order book replay, nanosecond-granularity microstructure, a volatility surface, execution-quality and slippage visualisation, or trillions of points where the question is not which indicator but whether the thing renders at all.
- Build custom when the chart is the product. If your differentiator is an AI-generated signal overlay, a stock rating rendered inline, or a sentiment band drawn against price, you want to own the render path rather than negotiate with someone else's plugin API for every visual idea.
- In practice the production systems we ship most often run both: the standard library for the trader's familiar workspace, and a WebGPU surface for the two or three views that are genuinely ours. They share the same data layer, which is the part that actually matters.
The AI Overlay Question
Since this is a chart in 2026, the signal overlay deserves a paragraph of its own, because it is where charting meets the rest of our work. Overlaying AI-generated buy and sell markers, ratings or sentiment bands directly on the price surface is straightforward to render - it is another instanced series - and deceptively easy to get wrong in a way that creates real liability. Two rules we hold to without exception.
- 01Signals are point-in-time artefacts and must be stored with the timestamp at which they were generated, not the timestamp of the bar they refer to. If you render a signal produced at 14:32 against the 14:30 bar with no distinction, you have built a look-ahead bias generator with a beautiful front end, and every backtest anyone runs against that view will be optimistic and wrong.
- 02Every rendered signal must carry provenance to the hover state: which model version, which inputs, generated when. This is the same decision-fingerprint discipline we apply to agents on the desk, and it is the only thing that makes an AI overlay defensible when someone asks why the chart told them to buy.
Practical Adoption Notes
- WebGPU needs a modern browser - Chrome and Edge 113+, with Firefox support having landed more recently. You need a Canvas2D or WebGL fallback path for the tail, and it should degrade to a lower point budget rather than to a blank panel.
- Test on integrated graphics, not just the workstation. A trading desk has good hardware; a client's laptop on a train does not, and the failure is graceful degradation only if you built it.
- Budget for the axis and label layer. It is easy to make the data render at 120fps and then spend your entire frame budget laying out tick labels in DOM. Render labels on a separate Canvas2D layer updated at a lower cadence - axes do not need to update at tick rate.
- Profile with the real feed at the real open. Synthetic uniform data hides every coalescing and reallocation bug you have, because the bugs live in burstiness.
- If you want a shortcut to evaluating whether this is worth it for you, clone ChartGPU, point its streaming candlestick example at your own feed, and watch what happens at your peak message rate. An afternoon of that is worth a month of architecture debate.
The Bottom Line
The browser charting ceiling that shaped a decade of trading front ends is gone, and the replacement is not a faster loop - it is a different division of labour. Move the reduction to a compute shader so downsampling happens where the data lives, collapse your render to one instanced draw per series so command submission stops dominating, push hit-testing onto the GPU so the crosshair stops fighting your main thread, and keep absolute time off the GPU so your intraday axis stays honest. Do those four things and a million ticks at 60fps is not a benchmark stunt, it is Tuesday. The reason this matters commercially is simple: when the chart can render everything, you stop designing around what the chart cannot show, and the analytics you have wanted to put in front of traders for years - full-depth replay, microstructure, execution quality, AI signals with provenance - become products rather than proposals. That is the work we do as a trading platform and charting developer in London, and this is the stack we would start from today.
References & Further Reading
- ChartGPU - open source WebGPU charting library (GitHub). github.com/ChartGPU/ChartGPU
- Show HN: ChartGPU - WebGPU-powered charting library (1M points at 60fps), January 2026. news.ycombinator.com/item?id=46706528
- WebGPU.com showcase - ChartGPU: WebGPU-powered charts that hold up at a million points. webgpu.com/showcase/chartgpu-webgpu-charts
- GIGAZINE - ChartGPU, the open-source WebGPU charting library that stays smooth at a million points. gigazine.net/gsc_news/en/20260122-chartgpu
- MDN - WebGPU API reference. developer.mozilla.org/en-US/docs/Web/API/WebGPU_API
- Can I Use - WebGPU browser support (for planning your fallback path). caniuse.com/webgpu
- Sveinn Steinarsson - Largest-Triangle-Three-Buckets reference implementation, with a link to the original thesis "Downsampling Time Series for Visual Representation". github.com/sveinn-steinarsson/flot-downsample
- W3C - WebGPU specification. w3.org/TR/webgpu
- W3C - WGSL: WebGPU Shading Language specification. w3.org/TR/WGSL
AlchmAI Engineering
Engineering, London
Written by the AlchmAI engineering team in Mayfair, London. We build trading platforms, real-time charts, market data pipelines and AI features for brokers, prop firms and fintech teams. The Playbook is where we explain how we approach these systems, with code you can run and sources you can check.
Code in this guide is illustrative and supplied without warranty. Review and test it before production use. Nothing here is investment advice. Important information