Hugging Face released SmolVLA this week, a 450M model with SO100/SO101 experiments. I'm tempted to try one very dull transfer rather than reproduce an impressive montage. For a first dataset, would you vary the starting positions straight away or establish one repeatable case first? I haven't trained anything yet.
https://huggingface.co/blog/smolvla
One case to check the recording pipeline, then vary positions. Keep some trials out of training from the start so you can tell whether the result transfers.