Skip to content

Adds Intel AMX TMUL hardware acceleration for FlashAttention using INT8 - #1049

Merged
copybara-service[bot] merged 1 commit into
devfrom
test_993708697
Oct 5, 2026
Merged

copybara-service[bot] merged 1 commit into
devfrom
test_993708697

Conversation

@copybara-service

@copybara-service copybara-service Bot commented Oct 5, 2026 •

Copy link
Copy Markdown

Adds Intel AMX TMUL hardware acceleration for FlashAttention using INT8
quantization:

  • Implements TileFlashAttentionAMX_Int8_Impl (TileFlashAttentionAMX_Int8_ImplT)
    on top of Highway's AMX-INT8 Tile64BMatMul wrappers (ComputeQKPinQ256AMX_Int8,
    ComputeQKTileAMX_Int8, SoftmaxAndQuantizeAMX_Int8, ComputePVGroupAMX_Int8,
    AccumulatePVTileInt8). The da/db signedness tags select TDPBSSD
    (int8 x int8) for Q*K^T and TDPBUSD (uint8 x int8) for P*V; tile
    config is compiler-managed via __tile1024i, and runtime detection uses
    hwy::HaveTile64BMatMulI8() (gated by GEMMA_HAVE_AMX_INT8).
  • Specializes qkv_dim == 256 (ComputeQKPinQ256AMX_Int8) by pinning a 16-query
    half-block of Q across tmm4..tmm7 (4 x 64 bytes) and streaming four
    32-token K tiles through tmm2..tmm3, cutting Q tile loads by 4x and
    making post-tilestored dequantization contiguous.
  • Vectorizes online softmax and 4-row VNNI packing (SoftmaxAndQuantizeAMX_Int8)
    with AVX-512 shuffle/permute intrinsics, and groups up to 4 query blocks
    (128 queries) per 128-token KV macro-step (ComputePVGroupAMX_Int8) with
    channel-outer / query-inner loop ordering to keep K, V, and P_vnni slices
    in the 48 KB L1D cache and the KV step slice in the 2 MB L2 cache.
  • Integrates AttentionImpl::kFlashAMXInt8 with quantized KV cache and query
    compression.
  • Adds INT8 AMX FlashAttention unit tests in flash_attention_test.cc and
    benchmarks.
  • Fixes a SIGILL on real AMX hardware in the BF16 kernel as well: the QK/PV
    helpers called Tile64BRelease(), but once inlined into the kernel loop the
    compiler does not reload the tile config after it, so the next iteration's
    tile ops raised #UD. Release now happens once after the loop.
  • Bumps the external Highway pin (MODULE.bazel, CMake) to
    353597727402dfc7b28e5b1474766e366ceafd24, which adds the AMX-INT8 wrappers.

@copybara-service
copybara-service Bot force-pushed the test_993708697 branch 3 times, most recently from 31def20 to 72df326 Compare October 5, 2026 22:23
quantization:
- Implements `TileFlashAttentionAMX_Int8_Impl` (`TileFlashAttentionAMX_Int8_ImplT`)
  on top of Highway's AMX-INT8 `Tile64BMatMul` wrappers (`ComputeQKPinQ256AMX_Int8`,
  `ComputeQKTileAMX_Int8`, `SoftmaxAndQuantizeAMX_Int8`, `ComputePVGroupAMX_Int8`,
  `AccumulatePVTileInt8`). The `da`/`db` signedness tags select `TDPBSSD`
  (`int8 x int8`) for `Q*K^T` and `TDPBUSD` (`uint8 x int8`) for `P*V`; tile
  config is compiler-managed via `__tile1024i`, and runtime detection uses
  `hwy::HaveTile64BMatMulI8()` (gated by `GEMMA_HAVE_AMX_INT8`).
- Specializes `qkv_dim == 256` (`ComputeQKPinQ256AMX_Int8`) by pinning a 16-query
  half-block of `Q` across `tmm4..tmm7` (`4 x 64` bytes) and streaming four
  32-token `K` tiles through `tmm2..tmm3`, cutting `Q` tile loads by 4x and
  making post-`tilestored` dequantization contiguous.
- Vectorizes online softmax and 4-row VNNI packing (`SoftmaxAndQuantizeAMX_Int8`)
  with AVX-512 shuffle/permute intrinsics, and groups up to 4 query blocks
  (128 queries) per 128-token KV macro-step (`ComputePVGroupAMX_Int8`) with
  channel-outer / query-inner loop ordering to keep `K`, `V`, and `P_vnni` slices
  in the 48 KB L1D cache and the KV step slice in the 2 MB L2 cache.
- Integrates `AttentionImpl::kFlashAMXInt8` with quantized KV cache and query
  compression.
- Adds INT8 AMX FlashAttention unit tests in `flash_attention_test.cc` and
  benchmarks.
- Fixes a SIGILL on real AMX hardware in the BF16 kernel as well: the QK/PV
  helpers called `Tile64BRelease()`, but once inlined into the kernel loop the
  compiler does not reload the tile config after it, so the next iteration's
  tile ops raised #UD. Release now happens once after the loop.
- Bumps the external Highway pin (MODULE.bazel, CMake) to
  353597727402dfc7b28e5b1474766e366ceafd24, which adds the AMX-INT8 wrappers.

PiperOrigin-RevId: 993941527
@copybara-service
copybara-service Bot merged commit 0224dcb into dev Oct 5, 2026
11 checks passed
@copybara-service
copybara-service Bot deleted the test_993708697 branch October 5, 2026 22:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant