주메뉴 바로가기 본문 바로가기

POPUP ZONE

2026 Processing-in-Memory Workshop 개최
회원 개인정보 추가 작성 요청
KAIST AI-PIM PIM반도체연구센터

IP 검색

CMSS (CXL Memory Sub-System)

CMSS is a Type-2 CXL device-side cache controller that maintains a hardware-coherent device cache over the CXL.cache and CXL.mem protocols. It allows an accelerator such as an NPU or GPU to cache host memory in device-attached DRAM close to computation while remaining fully coherent with the host CPU, removing the need for software-managed coherency. The IP is organized as three cooperating subsystems - a cache datapath (CMSS_CACHE_TOP), a device-attached memory controller (CMSS_MEM_TOP), and a CXL link controller (CMSS_CXL_TOP) - interconnected by credited CPI (CXL Protocol Interface) request, data, and response channels. An accelerator client attaches through a standard AXI4 slave interface, device DRAM attaches through two AXI4 master interfaces for the cache fill/eviction path and the CXL.mem path, configuration is provided through APB slave ports, and the host link is driven through a 4-lane PIPE interface to an external PCIe PHY.

Feature
· - Type-2 CXL device cache supporting CXL.cache (D2H/H2D) and CXL.mem (M2S/S2M) protocol flows
· - MESI coherency (Invalid / Shared / Exclusive / Modified) tracked per cache block
· - 64-byte cache block granularity with 62-bit cache tag and 64-bit ECC width
· - Device cache region backed by device DRAM with tag and data stored in memory; configurable capacity via CACHE_SIZE_LG2 (1 MB in the reference configuration)
· - Accelerator (NPU/GPU) client interface using standard AXI4 channels (AW/W/B/AR/R) with 64-bit address and 5-bit ID
· - Two separate AXI4 master interfaces for the cache fill/eviction path and the CXL.mem device-memory path, avoiding head-of-line blocking between the two flows
· - AxCACHE-driven caching policy selection per transaction, for example write-back write-allocate
· - 512-bit internal CPI datapath with credited request / data / response virtual channels
· - Slot-based data management with MSHR and AXI/CXL context buffers for multiple outstanding misses and in-order response return
· - CXL link layer with flit pack/unpack, CRC generation and checking, retry (replay) buffering, credit management, and TX arbitration
· - 4-lane PIPE interface toward an external PCIe PHY, with the PHY outside the IP boundary
· - APB slave configuration interfaces for the cache, memory, and CXL subsystems (32-bit address/data)
· - Three clock domains: aclk (AXI), cclk (CXL), and pclk (APB)
Application
· - CXL Type-2 accelerator devices such as NPUs, GPUs, and domain-specific accelerators that require c
Business Area
Data Center / AI & HPC - coherent interconnect and memory subsystem IP for CXL-attached accelerator (NPU/GPU) SoCs and CXL memory devices
Category

Interface Controller & PHY > CXL


Tech Specs
  • IP Name :

    CMSS (CXL Memory Sub-System)

  • Provider :

    Scalable Architecture Lab, Sungkyunkwan University (SKKU)

  • FPGA Device :

    Xilinx (AMD) Virtex UltraScale+ FPGA

  • Foundry :

    N/A

  • Bus Compliance :

    AMBA AXI4, AMBA APB

  • Technology :

    FPGA

  • Compliant Standard :

    Compute Express Link (CXL), PCIe PIPE interface

Deliverables
· - Synthesizable RTL source in SystemVerilog for CMSS_TOP and its three subsystems (CMSS_CACHE_TOP, CMSS_MEM_TOP, CMSS_CXL_TOP)
· - Parameter and interface packages (CMSS_PKG, CPI_PKG, PIPE_PKG) and interface definitions (AXI4_IF, APB_IF, PIPE_IF)
· - Digital datasheet describing architecture, port signal list, functional description, and cache state-transition reference
· - APB register specification for the cache, memory, and CXL subsystem configuration spaces
· - Simulation testbench and verification environment with directed coherency scenarios for CXL.cache and CXL.mem flows
· - Integration guide covering clock and reset domains, interface connection, and configuration sequence
· - Synthesis and FPGA implementation scripts with timing constraints
Validation Status
· Functionally verified by RTL simulation at both block level and top level. Block-level verification covers the cache datapath, the device-memory controller, and the CXL link controller individually; top-level verification exercises end-to-end CXL.cache and CXL.mem transaction flows between an AXI4 client, the device cache, and a host model. Coherency validation covers MESI state transitions for read and write transactions across AxCACHE policies and cache hit/miss combinations, including RdShared, RdOwnNoData, CleanEvictNoData, DirtyEvict, and the corresponding Go-S and Go-I host responses, as summarized in the cache state-transition reference of the datasheet. Link-layer functions including flit packing and unpacking, CRC generation and checking, retry buffer replay, and credit management have been verified in simulation. The design has been synthesized and validated on an FPGA prototyping platform for subsystem integration and coherency validation. Silicon validation has not been performed.
Availability
Not available for external distribution. This IP is developed and maintained internally by the Scalable Architecture Lab, Sungkyunkwan University, for research and academic purposes only, and is not offered for licensing, evaluation, or delivery to external parties at this time.
Functional Diagram
Benefits
· - Hardware-coherent device cache that lets an accelerator cache host memory without any software-managed coherency, reducing host-device data movement and programming complexity
· - Integrated CXL.mem controller so that device-attached memory and the coherent cache are served by a single IP, simplifying Type-2 device construction
· - Clean subsystem partitioning (cache datapath, memory controller, CXL link) connected by credited CPI channels, which decouples the cache and memory subsystems from CXL link timing and eases reuse or replacement of individual blocks
· - Slot-based data management with MSHR and context buffers, allowing multiple outstanding misses while returning responses in order to the client
· - Separate AXI4 master paths for cache fill/eviction and CXL.mem that avoid head-of-line blocking between the two traffic classes
· - Standard AXI4 and APB interfaces on all external ports, allowing drop-in integration into conventional AMBA-based SoC and FPGA platforms
· - Configurable cache capacity (CACHE_SIZE_LG2) and per-transaction caching policy selection through AxCACHE, enabling tuning without RTL changes
List