Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

You both seem to be talking past each other. There were a number of optimizations that made this possible. Some were with the model itself and are transferable, others are with the training pipeline and specific to the Nvidia hardware they trained on.


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: