系统仿真学报 ›› 2026, Vol. 38 ›› Issue (7): 2037-2052.doi: 10.16182/j.issn1004731x.joss.25-0847
张旭1,2, 刘珂1,2, 陈明宇1,2
收稿日期:2025-09-03
修回日期:2026-04-11
出版日期:2026-07-28
发布日期:2026-07-31
通讯作者:
刘珂
第一作者简介:张旭(1996-),男,助理研究员,博士,研究方向为计算机系统结构。
基金资助:Zhang Xu1,2, Liu Ke1,2, Chen Mingyu1,2
Received:2025-09-03
Revised:2026-04-11
Online:2026-07-28
Published:2026-07-31
Contact:
Liu Ke
摘要:
为解决现有CXL仿真平台在模拟大规模、CXL-以太网异构互连的数据中心场景时,难以兼顾仿真速度与拓扑灵活性的问题,提出了一种基于SoC-FPGA集群的半实物仿真平台CNetSim。采用SoC-FPGA硬件仿真终端节点,利用其真实的处理器与内存设备及FPGA实现CXL.mem协议,以支持标准Linux系统与真实分布式应用的高速运行;利用软件仿真灵活的拓扑配置能力,在仿真服务器上基于DPDK实现了CXL交换网络与机柜间以太网互连。仿真实验结果表明:相较于传统的软件模拟器,该平台在完整运行分布式应用时实现了至少3个数量级的加速,且能通过配置关键仿真参数精确等比例拟合真实系统的性能特征。
中图分类号:
张旭,刘珂,陈明宇 . 基于SoC-FPGA集群的CXL-以太网异构互连仿真平台[J]. 系统仿真学报, 2026, 38(7): 2037-2052.
Zhang Xu,Liu Ke,Chen Mingyu . Simulation Platform Based on SoC-FPGA Clusters for CXL-ethernet Heterogeneous Interconnection[J]. Journal of System Simulation, 2026, 38(7): 2037-2052.
| [1] | Malewicz G, Austern M H, Bik A J C, et al. Pregel: A System for Large-scale Graph Processing[C]//Proceedings of the 2010 ACM SIGMOD International Conference on Management of Data. New York: Association for Computing Machinery, 2010: 135-146. |
| [2] | Dean J, Ghemawat S. MapReduce: Simplified Data Processing on Large Clusters[J]. Communications of the ACM, 2008, 51(1): 107-113. |
| [3] | Liu Weibo, Wang Zidong, Liu Xiaohui, et al. A Survey of Deep Neural Network Architectures and Their Applications[J]. Neurocomputing, 2017, 234: 11-26. |
| [4] | Valiant L G. A Bridging Model for Parallel Computation[J]. Communications of the ACM, 1990, 33(8): 103-111. |
| [5] | Jia Zhihao, Kwon Y, Shipman Galen, et al. A Distributed Multi-GPU System for Fast Graph Processing[J]. Proceedings of the VLDB Endowment, 2017, 11(3): 297-310. |
| [6] | Jiang Yimin, Zhu Yibo, Lan Chang, et al. A Unified Architecture for Accelerating Distributed DNN Training in Heterogeneous GPU/CPU Clusters[C]//Proceedings of the 14th USENIX Conference on Operating Systems Design and Implementation. Berkeley: USENIX Association, 2020: 463-479. |
| [7] | Engelhardt Nina, Hayden Kwok-Hay So. GraVF: A Vertex-centric Distributed Graph Processing Framework on FPGAs[C]//2016 26th International Conference on Field Programmable Logic and Applications (FPL). Piscataway: IEEE, 2016: 1-4. |
| [8] | Li Huaicheng, Berger D S, Hsu L, et al. Pond: CXL-based Memory Pooling Systems for Cloud Platforms[C]//Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. New York: Association for Computing Machinery, 2023: 574-587. |
| [9] | Kwon Wonok, Park Chanho, Oh Myeonghoon. Gen-Z Memory Pool System Architecture[C]//2020 International Conference on Information and Communication Technology Convergence (ICTC). Piscataway: IEEE, 2020: 1356-1360. |
| [10] | Stuecheli J, Starke W J, Irish J D, et al. IBM POWER9 Opens up a New Era of Acceleration Enablement: OpenCAPI[J]. IBM Journal of Research and Development, 2018, 62(4/5): 8:1-8:8. |
| [11] | Tamimi Sajjad, Stock Florian, Koch Andreas, et al. An Evaluation of Using CCIX for Cache-coherent Host-FPGA Interfacing[C]//2022 IEEE 30th Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). Piscataway: IEEE, 2022: 1-9. |
| [12] | Das Sharma D, Blankenship R, Berger D. An Introduction to the Compute Express Link (CXL) Interconnect[J]. ACM Computing Surveys, 2024, 56(11): 290. |
| [13] | Sharma D D. PCI Express® 6.0 Specification at 64.0 GT/s with PAM-4 Signaling: A Low Latency, High Bandwidth, High Reliability and Cost-effective Interconnect[C]//2020 IEEE Symposium on High-Performance Interconnects (HOTI). Piscataway: IEEE, 2020: 1-8. |
| [14] | Guo Chuanxiong, Wu Haitao, Deng Zhong, et al. RDMA Over Commodity Ethernet at Scale[C]//Proceedings of the 2016 ACM SIGCOMM Conference. New York: Association for Computing Machinery, 2016: 202-215. |
| [15] | Sankar R. Memory Tiering, CXL, and RDMA Networking: Blending Advanced I/O Technologies to Get Past the AI Memory Wall[EB/OL]. [2025-09-02]. . |
| [16] | Ma Teng, Zhang Mingxing. Revisiting Distributed Memory in the CXL Era[EB/OL]. (2024-01-09) [2025-09-02]. |
| [17] | Rashinkar P, Paterson P, Singh L. System-on-a-chip Verification: Methodology and Techniques[M]. Boston, MA: Springer US, 2002. |
| [18] | Altera. What Is an SoC-FPGA?[EB/OL]. (2020-08-13) [2025-09-02]. . |
| [19] | Patel Harsh, Hiraskar Hrishikesh, Tahiliani Mohit P. Extending Network Emulation Support in Ns-3 Using DPDK[C]//Proceedings of the 2019 Workshop on Ns-3. New York: Association for Computing Machinery, 2019: 17-24. |
| [20] | Sasaki Kanon, Hirofuchi Takahiro, Yamaguchi Saneyasu, et al. An Accurate Packet Loss Emulation on a DPDK-based Network Emulator[C]//Proceedings of the 15th Asian Internet Engineering Conference. New York: Association for Computing Machinery, 2019: 1-8. |
| [21] | McVoy L W, Staelin C. Lmbench: Portable Tools for Performance Analysis[C]//USENIX Annual Technical Conference. Berkeley: USENIX Association, 1996: 279-294. |
| [22] | Tirumala A. Iperf: The TCP/UDP Bandwidth Measurement Tool[EB/OL]. [2025-09-02]. . |
| [23] | Schuh H N, Krishnamurthy A, Culler D, et al. CC-NIC: A Cache-coherent Interface to the NIC[C]//Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. New York: Association for Computing Machinery, 2024: 52-68. |
| [24] | Sano Shintaro, Bando Yosuke, Hiwada Kazuhiro, et al. GPU Graph Processing on CXL-based Microsecond-latency External Memory[C]//Proceedings of the SC '23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis. New York: Association for Computing Machinery, 2023: 962-972. |
| [25] | Abdullah R, Lee H, Zhou Huiyang, et al. Salus: Efficient Security Support for CXL-expanded GPU Memory[C]//2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA). Piscataway: IEEE, 2024: 1-15. |
| [26] | Arif M, Maurya A, Rafique M M. Accelerating Performance of GPU-based Workloads Using CXL[C]//Proceedings of the 13th Workshop on AI and Scientific Computing at Scale Using Flexible Computing. New York: Association for Computing Machinery, 2023: 27-31. |
| [27] | Gouk Donghyun, Kang Seungkwan, Bae Hanyeoreum, et al. Breaking Barriers: Expanding GPU Memory with Sub-two Digit Nanosecond Latency CXL Controller[C]//Proceedings of the 16th ACM Workshop on Hot Topics in Storage and File Systems. New York: Association for Computing Machinery, 2024: 108-115. |
| [28] | Fridman Yehonatan, Suprasad Mutalik Desai, Singh Navneet, et al. CXL Memory as Persistent Memory for Disaggregated HPC: A Practical Approach[C]//Proceedings of the SC '23 Workshops of the International Conference on High Performance Computing, Network, Storage, and Analysis. New York: Association for Computing Machinery, 2023: 983-994. |
| [29] | Sun Yan, Yuan Yifan, Yu Zeduo, et al. Demystifying CXL Memory with Genuine CXL-ready Systems and Devices[C]//Proceedings of the 56th Annual IEEE/ACM International Symposium on Microarchitecture. New York: Association for Computing Machinery, 2023: 105-121. |
| [30] | Cho A, Saxena A, Qureshi M, et al. A Case for CXL-centric Server Processors[EB/OL]. (2023-05-08) [2025-08-02]. . |
| [31] | Maruf H A, Wang Hao, Dhanotia A, et al. TPP: Transparent Page Placement for CXL-enabled Tiered-memory[C]//Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. New York: Association for Computing Machinery, 2023: 742-755. |
| [32] | Zhou Zhe, Xu Shuotao, Chen Yiqi, et al. Polaris: Enhancing CXL-based Memory Expanders with Memory-side Prefetching[C]//Advanced Parallel Processing Technologies. Singapore: Springer Nature Singapore, 2024: 19-39. |
| [33] | Yang Shaopeng, Kim Minjae, Nam Sanghyun, et al. Overcoming the Memory Wall with CXL-Enabled SSDs[C]//2023 USENIX Annual Technical Conference (USENIX ATC 23). Berkeley: USENIX Association, 2023: 601-617. |
| [34] | Zhou Zhe, Chen Yiqi, Zhang Tao, et al. NeoMem: Hardware/Software Co-design for CXL-native Memory Tiering[C]// Proceedings of the 57th Annual IEEE/ACM International Symposium on Microarchitecture. New York: Association for Computing Machinery, 2024: 1518-1531. |
| [35] | Kwon Miryeong, Lee Sangwon, Jung Myoungsoo. Cache in Hand: Expander-driven CXL Prefetcher for Next Generation CXL-SSD[C]//Proceedings of the 15th ACM Workshop on Hot Topics in Storage and File Systems. New York: Association for Computing Machinery, 2023: 24-30. |
| [36] | Jang Junhyeok, Choi Hanjin, Bae Hanyeoreum, et al. CXL-ANNS: Software-hardware Collaborative Memory Disaggregation and Computation for Billion-scale Approximate Nearest Neighbor Search[C]//2023 USENIX Annual Technical Conference (USENIX ATC 23). Berkeley: USENIX Association, 2023: 585-600. |
| [37] | Huangfu Wenqin, Malladi Krishna T, Chang Andrew, et al. BEACON: Scalable Near-data-processing Accelerators for Genome Analysis near Memory Pool with the CXL Support[C]//2022 55th IEEE/ACM International Symposium on Microarchitecture (MICRO). Piscataway: IEEE, 2022: 727-743. |
| [38] | Huo Pingyi, Devulapally A, Maruf H A, et al. PIFS-rec: Process-in-fabric-switch for Large-scale Recommendation System Inferences[C]//2024 57th IEEE/ACM International Symposium on Microarchitecture (MICRO). Piscataway: IEEE, 2024: 612-626. |
| [39] | Ha Minho, Ryu Junhee, Choi Jungmin, et al. Dynamic Capacity Service for Improving CXL Pooled Memory Efficiency[J]. IEEE Micro, 2023, 43(2): 39-47. |
| [40] | Lee Donghun, Willhalm Thomas, Ahn Minseon, et al. Elastic Use of Far Memory for In-memory Database Management Systems[C]//Proceedings of the 19th International Workshop on Data Management on New Hardware. New York: Association for Computing Machinery, 2023: 35-43. |
| [41] | Ahn Minseon, Willhalm Thomas, May Norman, et al. An Examination of CXL Memory Use Cases for In-memory Database Management Systems Using SAP HANA[J]. Proceedings of the VLDB Endowment, 2024, 17(12): 3827-3840. |
| [42] | Gouk Donghyun, Lee Sangwon, Kwon Miryeong, et al. Direct Access, High-performance Memory Disaggregation with DirectCXL[C]//2022 USENIX Annual Technical Conference (USENIX ATC 22). Berkeley: USENIX Association, 2022: 287-294. |
| [43] | Zhang Xu, Chang Yisong, Lu Tianyue, et al. Rethinking Design Paradigm of Graph Processing System with a CXL-like Memory Semantic Fabric[C]//2023 IEEE/ACM 23rd International Symposium on Cluster, Cloud and Internet Computing (CCGrid). Piscataway: IEEE, 2023: 25-35. |
| [44] | Miller E L, Benetopoulos A, Neville-Neil G V, et al. Pointers in Far Memory:A Rethink of How Data and Computations Should be Organized[J]. Queue, 2023, 21(3): 75-93. |
| [45] | Xu Yi, Mahar S, Liu Ziheng, et al. CXL Shared Memory Programming: Barely Distributed and Almost Persistent[EB/OL]. (2024-07-17) [2025-09-02]. . |
| [46] | Zhang Mingxing, Ma Teng, Hua Jinqi, et al. Partial Failure Resilient Memory Management System for (CXL-based) Distributed Shared Memory[C]//Proceedings of the 29th Symposium on Operating Systems Principles. New York: Association for Computing Machinery, 2023: 658-674. |
| [47] | Wang Chenjiu, He Ke, Fan Ruiqi, et al. CXL Over Ethernet: A Novel FPGA-based Memory Disaggregation Design in Data Centers[C]//2023 IEEE 31st Annual International Symposium on Field-Programmable Custom Computing Machines (FCCM). Piscataway: IEEE, 2023: 75-82. |
| [48] | Zhang Jie, Chen Xuzheng, Zhang Yin, et al. DmRPC: Disaggregated Memory-aware Datacenter RPC for Data-intensive Applications[C]//2024 IEEE 40th International Conference on Data Engineering (ICDE). Piscataway: IEEE, 2024: 3796-3809. |
| [49] | Wang Zhonghua, Guo Yixing, Lu Kai, et al. Rcmp: Reconstructing RDMA-based Memory Disaggregation via CXL[J]. ACM Transactions on Architecture and Code Optimization, 2024, 21(1): 15. |
| [50] | Tang Wenda, Han Ying, Ai Tianxiang, et al. Yggdrasil: Reducing Network I/O Tax with (CXL-based) Distributed Shared Memory[C]//Proceedings of the 53rd International Conference on Parallel Processing. New York: Association for Computing Machinery, 2024: 597-606. |
| [51] | Zhang Xu, Lu Tianyue, Chang Yisong, et al. Morpheus: An Adaptive DRAM Cache with Online Granularity Adjustment for Disaggregated Memory[C]//2023 IEEE 41st International Conference on Computer Design (ICCD). Piscataway: IEEE, 2023: 134-141. |
| [52] | Zhong Yijie, Zhou Minqiang, Shen Zhirong, et al. UniMem: Redesigning Disaggregated Memory Within a Unified Local-remote Memory Hierarchy[C]//Proceedings of the 2024 USENIX Conference on Usenix Annual Technical Conference. Berkeley: USENIX Association, 2024: 463-477. |
| [53] | Calciu I, Imran M T, Puddu Ivan, et al. Rethinking Software Runtimes for Disaggregated Memory[C]//Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems. New York: Association for Computing Machinery, 2021: 79-92. |
| [54] | Fang K, Peng D. NetDAM: Network Direct Attached Memory with Programmable In-memory Computing ISA[EB/OL]. (2021-10-28) [2025-09-02]. . |
| [55] | Zhong Yuhong, Berger D S, Waldspurger C, et al. Managing Memory Tiers with CXL in Virtualized Environments[C]//18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24). Berkeley: USENIX Association, 2024: 37-56. |
| [56] | Tang Yupeng, Zhou Ping, Zhang Wenhui, et al. Exploring Performance and Cost Optimization with ASIC-based CXL Memory[C]//Proceedings of the Nineteenth European Conference on Computer Systems. New York: Association for Computing Machinery, 2024: 818-833. |
| [57] | Puri Amit, Bellamkonda Kartheek, Narreddy Kailash, et al. DRackSim: Simulating CXL-enabled Large-scale Disaggregated Memory Systems[C]//Proceedings of the 38th ACM SIGSIM Conference on Principles of Advanced Discrete Simulation. New York: Association for Computing Machinery, 2024: 3-14. |
| [58] | Yang Yiwei, Safayenikoo P, Ma Jiacheng, et al. CXLMemSim: A Pure Software Simulated CXL.Mem for Performance Characterization[EB/OL]. (2023-03-10) [2025-09-02]. . |
| [59] | Wang Yanjing, Wu Lizhou, Hong Wentao, et al. CXL-DMSim: A Full-system CXL Disaggregated Memory Simulator with Comprehensive Silicon Validation[J]. IEEE Transactions on Computer-Aided Design of Integrated Circuits and Systems, 2026, 45(4): 1787-1801. |
| [60] | Boles D, Waddington D, Roberts D A. CXL-enabled Enhanced Memory Functions[J]. IEEE Micro, 2023, 43(2): 58-65. |
| [61] | Umeike J, Patel N, Manley A, et al. Profiling Gem5 Simulator[C]//2023 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). Piscataway: IEEE, 2023: 103-113. |
| [62] | 齐乐, 常轶松, 陈欲晓, 等. 基于SoC-FPGA的RISC-V处理器软硬件系统级平台[J]. 计算机研究与发展, 2023, 60(6): 1204-1215. |
| Qi Le, Chang Yisong, Chen Yuxiao, et al. A System-level Platform with SoC-FPGA for RISC-V Hardware-software Integration[J]. Journal of Computer Research and Development, 2023, 60(6): 1204-1215. | |
| [63] | JEDEC Solid State Technology Association. JEDEC Memory Module Reference Base Standard-for Compute Express Link (CXL)[EB/OL]. [2025-03-01]. . |
| [1] | 万士正, 程禹, 张旭, 范旭伟. 自适应视场角制导半实物仿真能力拓展方法[J]. 系统仿真学报, 2025, 37(3): 563-570. |
| [2] | 苏筱婷, 张小威, 田义, 李奇, 王帅豪. 星光导航动态仿真场景时序设计方法研究[J]. 系统仿真学报, 2025, 37(11): 2946-2955. |
| [3] | 李勇波, 田润梅, 张辉, 郭善鹏, 李琪. 基于Windows/RTX的实时仿测软件设计[J]. 系统仿真学报, 2024, 36(6): 1468-1474. |
| [4] | 豆建斌, 王小兵, 杨红坚, 高玉龙. 成像制导导弹试验鉴定半实物仿真系统设计与应用[J]. 系统仿真学报, 2024, 36(2): 522-532. |
| [5] | 刘鸿福, 付雅晶, 张万鹏, 张虎. 面向低资源的无人机指令意图识别算法及半实物仿真[J]. 系统仿真学报, 2024, 36(12): 2894-2905. |
| [6] | 常晓飞, 焦佳玥, 陈康, 符文星, 闫杰. 制导控制系统半实物仿真总体方案设计及集成联试[J]. 系统仿真学报, 2024, 36(1): 83-96. |
| [7] | 王天峥, 汤健, 夏恒, 乔俊飞. 城市固废焚烧过程的回路控制半实物仿真平台[J]. 系统仿真学报, 2023, 35(2): 241-253. |
| [8] | 刘紫寒, 侯凌霄, 李杨, 王智广, 张武龙. 基于自定义向导的通用实时半实物仿真代码自动生成方法[J]. 系统仿真学报, 2023, 35(10): 2279-2287. |
| [9] | 李启锐, 彭心怡. 基于深度强化学习的云作业调度及仿真研究[J]. 系统仿真学报, 2022, 34(2): 258-268. |
| [10] | 卢柏宏, 赵建军, 刘戈三. 电影虚拟摄影半实物仿真的研究与实现[J]. 系统仿真学报, 2021, 33(8): 1938-1946. |
| [11] | 严爱军, 夏恒, 刘溪芷. 城市生活垃圾焚烧过程监控半实物仿真平台研发[J]. 系统仿真学报, 2021, 33(6): 1427-1435. |
| [12] | 任福深, 孙雅琪, 胡庆, 李兆亮, 孙鹏宇, 范玉坤. 基于Unity3D的水下机器人半实物仿真系统[J]. 系统仿真学报, 2020, 32(8): 1546-1555. |
| [13] | 吴爱华, 赵不贿, 茅靖峰, 申海群, 张旭东. 基于快速控制原型的风力发电半实物仿真系统[J]. 系统仿真学报, 2020, 32(3): 482-491. |
| [14] | 张翔, 刘梦焱, 张鹏, 朱克炜, 闫小东. 一种激光驾束制导武器半实物仿真系统构建及试验方法[J]. 系统仿真学报, 2020, 32(11): 2192-2198. |
| [15] | 王林鹏, 王超磊, 戴玉婷. 转台外置的空馈式低频寻的制导半实物仿真研究[J]. 系统仿真学报, 2020, 32(1): 69-77. |
| 阅读次数 | ||||||
|
全文 |
|
|||||
|
摘要 |
|
|||||