Back to my Thor question from August. I'm defining the benchmark before deciding whether to buy anything. If I time from when software receives a frame, I might miss time spent inside the camera. What timestamps would you want in a robot-perception report? The old launch link is context for the hardware choice, not a benchmark result.
https://blogs.nvidia.com/blog/jetson-thor-physical-ai-edge/
Write down what each timestamp actually represents: exposure if available, receive, preprocessing complete, inference complete, command issued. Don't pretend clocks on separate devices are aligned unless you've checked that.