Research & Papers

DxPTA: Photonic Transformer Accelerators hit 15.2x faster design search

Photonic chips for AI achieve 6ms latency under strict power and area constraints

Deep Dive

Transformer-based AI models are approaching AGI-level performance, but their massive size makes efficient hardware implementation a bottleneck. Recent photonic transformer accelerators (PTAs) offer dramatic speed and energy gains over electronic chips, yet existing PTA designs are hand-tuned and ignore application constraints like area, power, and latency. This manual approach is unscalable and time-consuming.

To solve this, Rachmad Vidya Wicaksana Putra and colleagues from NYU Abu Dhabi introduce DxPTA — a methodology that systematically explores the PTA design space using a coherent optical dataflow-guided strategy. It identifies critical architecture parameters, analyzes their impact, and runs a constraint-aware search algorithm. For models like DeiT-T/S/B and BERT-B/L, DxPTA produces designs meeting tight constraints (50mm², 5W, 50mJ, 10ms) while consuming as little as 26mm², 4.8W, 39mJ, and 6ms. The search is 15.2× faster than exhaustive enumeration, making photonic AI accelerators practical for real-world deployment.

Key Points
  • DxPTA automates HW/SW co-design for photonic transformer accelerators using optical dataflow analysis
  • Achieves designs meeting 50mm² area, 5W power, 50mJ energy, and 10ms latency constraints for DeiT and BERT models
  • 15.2× faster design space search than exhaustive methods, enabling scalable PTA development

Why It Matters

Photonic accelerators could drastically cut AI energy costs; DxPTA makes them designable within real-world constraints.

📬 Get the top 10 AI stories daily