KESTREL
OverviewInternalsThe Die DatasheetDevelopersOrder

Kestrel Silicon · Falco architecture · Tape-out 2

Talon
X90

The first workstation GPU designed from the desk up — 48 GB of GDDR7, 1.15 TB/s, 285 watts, and a datasheet we'll show anyone.

The honest datasheet

Q1 2026 · $3,499 · fully refundable deposit · 4-year warranty

76.5TFLOPS FP32
1.15TB/s bandwidth
285watts board
FIG. 01 — ASSEMBLED · 1:1 DRAG TO ORBIT KS-102 · 48 GB GDDR7 ECC 328 × 140 × 61 mm

01 — Overview

Not a gaming chip
in a blazer.

Workstation GPUs have spent fifteen years as gaming silicon with ECC bolted on and a price with feelings involved. Talon X90 starts from a different question: what does the desk actually need? Not a bigger number — bandwidth you can saturate, memory you can trust, a card you can open, and software that gets out of the way.

0TFLOPS

FP32, sustained all-core burst at 2.62 GHz

0TB/s

384-bit · 12 × 32 Gb GDDR7, ECC on-die

0GB

VRAM, error-corrected — no silent flips

0W

board power — cooler than the card it replaces

0B

transistors · TSMC 4NP · 524 mm²

$0

launch price. Flagship-class silicon without the flagship tax.

02 — Internals

Open it up.

Every professional card should survive a curious owner. We made the X90 explainable: seven layers, one slider, no glue. Drag the model too — it doesn't mind.

FIG. 02 — EXPLODED VIEW DRAG — ORBIT · SCROLL — ZOOM
0 — Assembled Exploded — 100
    / 01

    Vapor chamber

    One billet, nickel-plated copper. 44 fins at 0.4 mm and six 8 mm sintered-wick heatpipes. Rated for 285 W at 1,900 m.

    / 02

    Triple 92 mm fans

    Nine blades each, fluid-dynamic bearings, 32.4 dBA at full tilt. Zero RPM below 50 °C — your studio stays quieter than the fridge.

    / 03

    16-phase VRM

    DrMOS stages rated 540 A on a 14-layer PCB with 2 oz copper. Transients from 0 to full load without a flicker.

    / 04

    No RGB.

    One warm-white status LED on the top rail that dims with the fans. It's a tool. The card is held together with Torx and opinions.

    03 — The die

    58.9 billion transistors,
    drawn to scale.

    Floorplan of the KS-102, our second tape-out. Hover anything — every block is labeled, because a die shot without annotations is just modern art.

    L2 — 48 MB (2 × 24 MB) MEDIA ×2 — AV1 DISPLAY — DP 2.1 PCIE 5.0 GSP KS-102 · 524 mm² · TSMC 4NP

    04 — The honest datasheet

    We win most rows.
    Not all. Read it.

    We'll spare you the spider chart. Competitor numbers are quoted straight from vendor spec sheets, unedited. Ours are measured on preproduction silicon. Rust marks the best value in each row — including the rows we don't own.

    Specification Talon X902026 · Kestrel RTX A60002020 · NVIDIA RTX 6000 Ada2022 · NVIDIA Radeon PRO W79002023 · AMD GeForce RTX 40902022 · NVIDIA
    Architecture / dieFalco · KS-102Ampere · GA102Ada · AD102RDNA 3 · Navi 31Ada · AD102
    ProcessTSMC 4NPSamsung 8NTSMC 4NTSMC N5 + N6TSMC 4N
    Transistors58.9 B28.3 B76.3 B57.7 B76.3 B
    Shader cores14,59210,75218,1766,14416,384
    Boost clock2.62 GHz1.80 GHz2.505 GHz2.50 GHz2.52 GHz
    FP32 compute76.5 TFLOPS38.7 TFLOPS91.1 TFLOPS61 TFLOPS82.6 TFLOPS
    Ray-tracing units1148414296128
    Matrix / AI units228336568192512
    Memory48 GB GDDR7 ECC48 GB GDDR6 ECC48 GB GDDR6 ECC48 GB GDDR6 ECC24 GB GDDR6X
    Memory bandwidth1,152 GB/s768 GB/s960 GB/s864 GB/s1,008 GB/s
    FP641.20 TFLOPS1.21 TFLOPS1.42 TFLOPS3.82 TFLOPS1.29 TFLOPS
    Board power285 W300 W300 W295 W450 W
    Host interfacePCIe 5.0 ×16PCIe 4.0 ×16PCIe 4.0 ×16PCIe 4.0 ×16PCIe 4.0 ×16
    ECC memoryYesYesYesYesNo
    Display outputs4× DP 2.1 + HDMI 2.1b4× DP 1.4a4× DP 1.4a3× DP 2.1 + HDMI3× DP 1.4a + HDMI
    Warranty4 years3 years3 years3 years1–3 years⁴
    Launch MSRP$3,499$4,650$6,800$3,999$1,599
    FP32 per $1,00021.9 TFLOPS8.3 TFLOPS13.4 TFLOPS15.3 TFLOPS51.7 TFLOPS

    best in row  ·  The X90 takes seven rows outright, ties two, and cedes the silicon-count rows to a chip that costs twice as much. Then there's row seventeen.

    Where the bandwidth actually shows

    Internal bench index · RTX 6000 Ada = 100 · higher is better · preproduction units⁵

    Viewport — 12 M-triangle assembly, shaded

    Talon X9095
    RTX A600048
    RTX 6000 Ada100
    W790069
    RTX 409098

    Path tracing — USD studio scene, Cycles-class

    Talon X9088
    RTX A600030
    RTX 6000 Ada100
    W790042
    RTX 409091

    Fluid sim — 40 M-particle SPH, bandwidth-bound

    Talon X90126
    RTX A600050
    RTX 6000 Ada100
    W790079
    RTX 409097

    LLM LoRA fine-tune — 7B, FP8 accumulate

    Talon X90108
    RTX A600030
    RTX 6000 Ada100
    W790056
    RTX 409047

    The pattern is the point. With 114 RT units against NVIDIA's 142, raw path-tracing throughput is physics, not marketing — we lose that row and say so. But scenes that thrash 24 GB cards run resident here, and bandwidth-bound simulation is where the extra 192 GB/s earns its keep.

    1Competitor figures from NVIDIA and AMD published product specifications for RTX A6000, RTX 6000 Ada Generation, GeForce RTX 4090 and Radeon PRO W7900 (retrieved September 2025). Launch MSRPs as announced.

    2Talon X90 figures: preproduction boards, Kestrel labs, 23 °C ambient, stock 285 W power limit.

    3FP64 rates: 1/64 of FP32 on Falco and Ada; 1/16 of FP32 on RDNA 3.

    4GeForce RTX 4090 warranty varies by board partner, typically 1–3 years.

    5Bench index: internal harness covering viewport shading, path tracing, SPH simulation and LoRA fine-tuning. The harness ships with the Fledge SDK so you can reproduce — or dispute — every bar above.

    05 — Software

    A GPU you can
    actually program.

    Fledge is our SDK, and it has one rule: the common thing should be the easy thing. One installer, real error messages, and a first render before your coffee cools.

    • Errors are exceptions — not 47 return codes to grep
    • printf() inside kernels, real breakpoints in VS Code
    • SPIR-V native — GLSL, HLSL, WGSL and Slang all compile
    • CUDA Bridge ports existing projects: 91.6% auto, rest flagged by line
    • Compiler cache — second build is 40× faster
    Fledge CPythonRustVulkan 1.4OpenGL 4.6OpenCL 3.0DirectX 12 UltimateWebGPUCUDA Bridge
    5.7× fewer lines, median, across 200 CUDA samples ported by our team. Counted, not vibes.
    // first_flight.c — the whole program
    #include <fledge.h>
    
    kernel matmul(tensor<f32> A, tensor<f32> B,
                  tensor<f32> out C) {
      f32 acc = 0;
      for (u32 k = 0; k < A.cols; ++k)
        acc += A[row, k] * B[k, col];
      C[row, col] = acc;
    }
    
    int main() {
      auto dev = fledge::device(0);   // Talon X90
      auto A = dev.load("A.f32"), B = dev.load("B.f32");
      auto C = dev.zeros(A.rows, B.cols);
      dev.fly(matmul, A, B, C);        // schedule + sync. that's it.
      C.save("C.f32");
    }
    # pip install fledge — that's the whole install
    import fledge as fl
    
    dev = fl.devices()[0]              # Talon X90 (48 GB)
    a = fl.randn(dev, 8192, 8192, dtype=fl.f32)
    b = fl.randn(dev, 8192, 8192, dtype=fl.f32)
    
    c = a @ b                            # matrix cores, FP8 accumulate
    print(f"{c.sum():.1f}")             # ~1.9 ms later
    # bring your existing CUDA project
    $ fledge port ./train.cu
      scanning … 214 kernel launches, 38 cuda* calls
      auto-ported 196/214 (91.6%) · 18 flagged with line hints
      wrote ./train.fg · conformance: PASS (7/7 tests)
    
    $ fledge run ./train.fg --device talonx90
      epoch 1/3  loss 2.317  4.6 it/s
      # (RTX 6000 Ada baseline on same harness: 4.1 it/s)
    $ python quickstart.py
    # press RUN — output replayed from our lab, Nov 2025
    talonx90 · fledge 1.4 · tty0

    The ceremony tax

    The same 4096² GEMM, host side. Every allocation, copy and error check in the CUDA version is real code that real people maintain. Fledge makes the safe path the short path.

    Typical CUDA host code — 16 lines

    1cudaMalloc(&dA, N*N*4); cudaMalloc(&dB, N*N*4);
    2cudaMalloc(&dC, N*N*4);
    3cudaMemcpy(dA, hA, N*N*4, cudaMemcpyHostToDevice);
    4cudaMemcpy(dB, hB, N*N*4, cudaMemcpyHostToDevice);
    5dim3 blk(16,16), grd((N+15)/16,(N+15)/16);
    6matmul<<<grd, blk>>>(dA, dB, dC, N);
    7cudaError_t e = cudaGetLastError();
    8if (e != cudaSuccess) {
    9  fprintf(stderr, "%s\n", cudaGetErrorString(e));
    10  return 1;
    11}
    12cudaMemcpy(hC, dC, N*N*4, cudaMemcpyDeviceToHost);
    13cudaFree(dA); cudaFree(dB); cudaFree(dC);
    14// …and pray the launch config fits the SM count

    Fledge — 5 lines

    1auto A = dev.load("A.f32");
    2auto B = dev.load("B.f32");
    3auto C = dev.zeros(N, N);
    4dev.fly(matmul, A, B, C);
    5C.save("C.f32");
    6// copies, tiling and error checks are the runtime's job

    06 — Order

    $3,499.
    No asterisk.

    • Ships Q1 2026 — 2,000 units in the first run, built in Tampere
    • 24-hour burn-in at 285 W before every card leaves the floor
    • 4-year warranty · 30-day no-questions returns
    • Fan, cable and bracket kits sold for a minimum of 7 years
    Read the SDK docs
    328 mm 140 mm 61 mm 3 × 92 mm VORTEX FANS · ZERO-RPM < 50°C 4 × DP 2.1 UHBR20 + HDMI 2.1b · 12V-2×6 (ADAPTER INCL.) 3-SLOT 12V-2×6
    FIG. 03 — Production drawingREV C · SHEET 1/1