DxPTA: Photonic Transformer Accelerators hit 15.2x faster design search
Photonic chips for AI achieve 6ms latency under strict power and area constraints
Transformer-based AI models are approaching AGI-level performance, but their massive size makes efficient hardware implementation a bottleneck. Recent photonic transformer accelerators (PTAs) offer dramatic speed and energy gains over electronic chips, yet existing PTA designs are hand-tuned and ignore application constraints like area, power, and latency. This manual approach is unscalable and time-consuming.
To solve this, Rachmad Vidya Wicaksana Putra and colleagues from NYU Abu Dhabi introduce DxPTA — a methodology that systematically explores the PTA design space using a coherent optical dataflow-guided strategy. It identifies critical architecture parameters, analyzes their impact, and runs a constraint-aware search algorithm. For models like DeiT-T/S/B and BERT-B/L, DxPTA produces designs meeting tight constraints (50mm², 5W, 50mJ, 10ms) while consuming as little as 26mm², 4.8W, 39mJ, and 6ms. The search is 15.2× faster than exhaustive enumeration, making photonic AI accelerators practical for real-world deployment.
- DxPTA automates HW/SW co-design for photonic transformer accelerators using optical dataflow analysis
- Achieves designs meeting 50mm² area, 5W power, 50mJ energy, and 10ms latency constraints for DeiT and BERT models
- 15.2× faster design space search than exhaustive methods, enabling scalable PTA development
Why It Matters
Photonic accelerators could drastically cut AI energy costs; DxPTA makes them designable within real-world constraints.