SE Shared IDE For PCIe Bifurcation: Scaling Security Without Scaling Complexity & Resources
Posted: Thu Aug 06, 2026 7:06 am
PCIe is the primary high-speed interconnect used to connect processors with accelerators, memory devices, storage, and networking components. It provides a common and scalable foundation for moving large volumes of data between the compute and I/O resources that make up a modern system. As architectures add more powerful accelerators and higher-performance devices, PCIe continues to evolve to deliver greater bandwidth while maintaining low latency and broad ecosystem interoperability. These capabilities have made it essential to data center, AI, cloud, networking, storage, and other performance-intensive systems. PCIe security is no longer optional As PCIe links carry increasingly sensitive workloads, protecting data in motion has become as important as bandwidth, latency, and interoperability. Data moving between processors, accelerators, memory devices, storage subsystems, and network adapters must be protected against interception, tampering, and replay attacks, particularly in AI training, confidential computing, cloud infrastructure, and high-performance data processing. The PCIe specification defines Integrity and Data Encryption (IDE) as the standard for securing data across a PCIe link. IDE protects Transaction Layer Packets with confidentiality, integrity, and replay protection as they travel between connected components. Bifurcation is becoming more common A growing trend to support more devices and make better use of available lanes is PCIe bifurcation, which divides a single wider interface into multiple independent links. For example, a PCIe x16 interface can be configured as two x8 links, four x4 links, or a combination such as x8 + x4 + x2 + x2. The total lane count remains unchanged, but those lanes form more independent connections. Bifurcation gives system designers greater flexibility, improves lane utilization, supports multiple endpoints, and enables reduced resource usage. That flexibility also changes the security model: one x16 link is not equivalent to four separate links because each independent connection requires its own IDE protection, as illustrated in the figure below. This creates an implementation challenge for traditional architectures, which typically dedicate IDE resources to each PCIe controller or port. As the number of independent links grows, designers may need additional encryption engines and management logic, increasing silicon area, power, and design complexity.
Why traditional IDE architectures become inefficient With bifurcation increasing the number of links, designers often find themselves replicating encryption engines, management logic, and supporting security infrastructure. This approach introduces several challenges, such as:
The solution combines high-throughput AES-GCM and SM4-GCM processing with a Shared IDE wrapper, segmented bus architecture, dynamic bandwidth allocation, and support for bifurcated PCIe configurations. Its centralized management and configurable pipeline stages simplify integration while preserving the integrity, isolation, and standards alignment required by PCIe IDE. The main benefits of the solution include:
Source: https://semiengineering.com/shared-ide- ... resources/
Why traditional IDE architectures become inefficient With bifurcation increasing the number of links, designers often find themselves replicating encryption engines, management logic, and supporting security infrastructure. This approach introduces several challenges, such as: - Area growth: Each additional protected link requires additional hardware resources. As the number of links scales, area consumption can grow significantly.
- Floorplanning complexity: Security logic must often be placed close to individual controllers, creating congestion and making physical implementation more difficult.
- Scalability limitations: Security resources scale in proportion to controller count rather than actual throughput requirements. This can produce inefficiencies in highly bifurcated designs.
- Lower cost and area: By eliminating unnecessary duplication, Shared IDE can reduce overall hardware requirements. This leads to fewer replicated security blocks, lower silicon area, and improved power efficiency.
- Greater design flexibility: Shared IDE can be physically separated from PCIe controllers and PHYs, giving implementation teams more freedom during floorplanning, which helps to reduce congestion and facilitates easier timing closure.
- Security without compromise: Importantly, sharing resources does not mean sharing security contexts. Each port retains independent protection, dedicated security contexts, separate key management, and standards-aligned PCIe IDE compliance. This means security remains isolated even while implementation resources are shared.
The solution combines high-throughput AES-GCM and SM4-GCM processing with a Shared IDE wrapper, segmented bus architecture, dynamic bandwidth allocation, and support for bifurcated PCIe configurations. Its centralized management and configurable pipeline stages simplify integration while preserving the integrity, isolation, and standards alignment required by PCIe IDE. The main benefits of the solution include: - Improved PPA: Shared processing reduces duplicated security hardware, with potential gate-count reductions of up to ~1.5x compared with traditional approaches.
- Better floorplanning: Security blocks can be placed independently of controllers and PHYs, reducing congestion in critical design regions.
- Bifurcation-friendly scaling: Security resource requirements do not need to increase in direct proportion to the number of protected links.
- Flexible, standards-aligned integration: Configurable deployment options support different architectures while maintaining independent protection for each port.
Source: https://semiengineering.com/shared-ide- ... resources/