Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Sure, the "6b subset" can be more knowledgeable on its area than a whole 27b generalist (and more efficient), but where is the simulated Intelligence encoded? A 6b subset as or more intelligent than a 27b raises the question of how metacognition skills are stored.


MOEs are built by training a second "router" model to identify which parts matter inside the dense model.

Think of MOE as ignoring noise, rather than a more efficient encoding of the data we teach it which ones can be ignored. Turning down the noise actually sharpens the results sometimes.

Modern LLM's are wildly inefficient.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: