Coming soon — deep learning model compilation, low-latency serving, hardware utilization, and inference engine optimizations.