Kestrel Silicon · Falco architecture · Tape-out 2
The first workstation GPU designed from the desk up — 48 GB of GDDR7, 1.15 TB/s, 285 watts, and a datasheet we'll show anyone.
Q1 2026 · $3,499 · fully refundable deposit · 4-year warranty
01 — Overview
Workstation GPUs have spent fifteen years as gaming silicon with ECC bolted on and a price with feelings involved. Talon X90 starts from a different question: what does the desk actually need? Not a bigger number — bandwidth you can saturate, memory you can trust, a card you can open, and software that gets out of the way.
FP32, sustained all-core burst at 2.62 GHz
384-bit · 12 × 32 Gb GDDR7, ECC on-die
VRAM, error-corrected — no silent flips
board power — cooler than the card it replaces
transistors · TSMC 4NP · 524 mm²
launch price. Flagship-class silicon without the flagship tax.
Our first in-house design. 96 Falcon clusters with a burst scheduler — the kestrel hunts in short strikes, so does the X90: 2.62 GHz when the workload asks, 0 W when it doesn't.
48 GB of GDDR7 ECC on a 384-bit bus. Scenes that swap on 24 GB cards run resident here — and 1.15 TB/s means the swap file stays bored.
Torx, not glue. The shroud, fans and backplate come off with six screws. Fan and cable kits will be sold separately for seven years, minimum.
One installer, three languages, real error messages. Port existing CUDA source with the built-in Bridge — measured at 91.6% auto-conversion.
02 — Internals
Every professional card should survive a curious owner. We made the X90 explainable: seven layers, one slider, no glue. Drag the model too — it doesn't mind.
One billet, nickel-plated copper. 44 fins at 0.4 mm and six 8 mm sintered-wick heatpipes. Rated for 285 W at 1,900 m.
Nine blades each, fluid-dynamic bearings, 32.4 dBA at full tilt. Zero RPM below 50 °C — your studio stays quieter than the fridge.
DrMOS stages rated 540 A on a 14-layer PCB with 2 oz copper. Transients from 0 to full load without a flicker.
One warm-white status LED on the top rail that dims with the fans. It's a tool. The card is held together with Torx and opinions.
03 — The die
Floorplan of the KS-102, our second tape-out. Hover anything — every block is labeled, because a die shot without annotations is just modern art.
04 — The honest datasheet
We'll spare you the spider chart. Competitor numbers are quoted straight from vendor spec sheets, unedited. Ours are measured on preproduction silicon. Rust marks the best value in each row — including the rows we don't own.
| Specification | Talon X902026 · Kestrel | RTX A60002020 · NVIDIA | RTX 6000 Ada2022 · NVIDIA | Radeon PRO W79002023 · AMD | GeForce RTX 40902022 · NVIDIA |
|---|---|---|---|---|---|
| Architecture / die | Falco · KS-102 | Ampere · GA102 | Ada · AD102 | RDNA 3 · Navi 31 | Ada · AD102 |
| Process | TSMC 4NP | Samsung 8N | TSMC 4N | TSMC N5 + N6 | TSMC 4N |
| Transistors | 58.9 B | 28.3 B | 76.3 B | 57.7 B | 76.3 B |
| Shader cores | 14,592 | 10,752 | 18,176 | 6,144 | 16,384 |
| Boost clock | 2.62 GHz | 1.80 GHz | 2.505 GHz | 2.50 GHz | 2.52 GHz |
| FP32 compute | 76.5 TFLOPS | 38.7 TFLOPS | 91.1 TFLOPS | 61 TFLOPS | 82.6 TFLOPS |
| Ray-tracing units | 114 | 84 | 142 | 96 | 128 |
| Matrix / AI units | 228 | 336 | 568 | 192 | 512 |
| Memory | 48 GB GDDR7 ECC | 48 GB GDDR6 ECC | 48 GB GDDR6 ECC | 48 GB GDDR6 ECC | 24 GB GDDR6X |
| Memory bandwidth | 1,152 GB/s | 768 GB/s | 960 GB/s | 864 GB/s | 1,008 GB/s |
| FP64 | 1.20 TFLOPS | 1.21 TFLOPS | 1.42 TFLOPS | 3.82 TFLOPS | 1.29 TFLOPS |
| Board power | 285 W | 300 W | 300 W | 295 W | 450 W |
| Host interface | PCIe 5.0 ×16 | PCIe 4.0 ×16 | PCIe 4.0 ×16 | PCIe 4.0 ×16 | PCIe 4.0 ×16 |
| ECC memory | Yes | Yes | Yes | Yes | No |
| Display outputs | 4× DP 2.1 + HDMI 2.1b | 4× DP 1.4a | 4× DP 1.4a | 3× DP 2.1 + HDMI | 3× DP 1.4a + HDMI |
| Warranty | 4 years | 3 years | 3 years | 3 years | 1–3 years⁴ |
| Launch MSRP | $3,499 | $4,650 | $6,800 | $3,999 | $1,599 |
| FP32 per $1,000 | 21.9 TFLOPS | 8.3 TFLOPS | 13.4 TFLOPS | 15.3 TFLOPS | 51.7 TFLOPS |
best in row · The X90 takes seven rows outright, ties two, and cedes the silicon-count rows to a chip that costs twice as much. Then there's row seventeen.
Internal bench index · RTX 6000 Ada = 100 · higher is better · preproduction units⁵
The pattern is the point. With 114 RT units against NVIDIA's 142, raw path-tracing throughput is physics, not marketing — we lose that row and say so. But scenes that thrash 24 GB cards run resident here, and bandwidth-bound simulation is where the extra 192 GB/s earns its keep.
1Competitor figures from NVIDIA and AMD published product specifications for RTX A6000, RTX 6000 Ada Generation, GeForce RTX 4090 and Radeon PRO W7900 (retrieved September 2025). Launch MSRPs as announced.
2Talon X90 figures: preproduction boards, Kestrel labs, 23 °C ambient, stock 285 W power limit.
3FP64 rates: 1/64 of FP32 on Falco and Ada; 1/16 of FP32 on RDNA 3.
4GeForce RTX 4090 warranty varies by board partner, typically 1–3 years.
5Bench index: internal harness covering viewport shading, path tracing, SPH simulation and LoRA fine-tuning. The harness ships with the Fledge SDK so you can reproduce — or dispute — every bar above.
05 — Software
Fledge is our SDK, and it has one rule: the common thing should be the easy thing. One installer, real error messages, and a first render before your coffee cools.
// first_flight.c — the whole program #include <fledge.h> kernel matmul(tensor<f32> A, tensor<f32> B, tensor<f32> out C) { f32 acc = 0; for (u32 k = 0; k < A.cols; ++k) acc += A[row, k] * B[k, col]; C[row, col] = acc; } int main() { auto dev = fledge::device(0); // Talon X90 auto A = dev.load("A.f32"), B = dev.load("B.f32"); auto C = dev.zeros(A.rows, B.cols); dev.fly(matmul, A, B, C); // schedule + sync. that's it. C.save("C.f32"); }
# pip install fledge — that's the whole install import fledge as fl dev = fl.devices()[0] # Talon X90 (48 GB) a = fl.randn(dev, 8192, 8192, dtype=fl.f32) b = fl.randn(dev, 8192, 8192, dtype=fl.f32) c = a @ b # matrix cores, FP8 accumulate print(f"{c.sum():.1f}") # ~1.9 ms later
# bring your existing CUDA project $ fledge port ./train.cu scanning … 214 kernel launches, 38 cuda* calls auto-ported 196/214 (91.6%) · 18 flagged with line hints wrote ./train.fg · conformance: PASS (7/7 tests) $ fledge run ./train.fg --device talonx90 epoch 1/3 loss 2.317 4.6 it/s # (RTX 6000 Ada baseline on same harness: 4.1 it/s)
$ python quickstart.py # press RUN — output replayed from our lab, Nov 2025
The same 4096² GEMM, host side. Every allocation, copy and error check in the CUDA version is real code that real people maintain. Fledge makes the safe path the short path.
1cudaMalloc(&dA, N*N*4); cudaMalloc(&dB, N*N*4); 2cudaMalloc(&dC, N*N*4); 3cudaMemcpy(dA, hA, N*N*4, cudaMemcpyHostToDevice); 4cudaMemcpy(dB, hB, N*N*4, cudaMemcpyHostToDevice); 5dim3 blk(16,16), grd((N+15)/16,(N+15)/16); 6matmul<<<grd, blk>>>(dA, dB, dC, N); 7cudaError_t e = cudaGetLastError(); 8if (e != cudaSuccess) { 9 fprintf(stderr, "%s\n", cudaGetErrorString(e)); 10 return 1; 11} 12cudaMemcpy(hC, dC, N*N*4, cudaMemcpyDeviceToHost); 13cudaFree(dA); cudaFree(dB); cudaFree(dC); 14// …and pray the launch config fits the SM count
1auto A = dev.load("A.f32"); 2auto B = dev.load("B.f32"); 3auto C = dev.zeros(N, N); 4dev.fly(matmul, A, B, C); 5C.save("C.f32"); 6// copies, tiling and error checks are the runtime's job
06 — Order
$3,499.
No asterisk.
FULLY REFUNDABLE $99 DEPOSIT · LOCKS LAUNCH PRICING · SHIPS Q1 2026
We'll email you when your build slot opens. Deposits are fully refundable until shipment.