Wider AI Models Can Now Learn Step by Step — and It's Cheaper
This could cut the cost and energy needed to build tomorrow's AI.
In self-supervised learning, models can be trained greedily one layer at a time — no end-to-end backpropagation of error — and still learn comparable representations, as long as the network is made wider. Researchers investigated how width and depth affect greedy layer-wise versus end-to-end self-supervised training in convolutional networks, and found that in wider networks the benefits of end-to-end backpropagation over greedy layer-wise training shrink. In relatively shallow and very wide networks, they even observed higher performance in models trained with greedy layer-wise training. Analysis of the representations showed that very wide greedy-trained networks exhibit more favorable representational geometry than networks trained end-to-end. The work shows that width can compensate for restricted credit assignment, and identifies differences in representational geometry as a potential mechanism for the improved performance.
- AI can be taught one layer at a time instead of all at once, if the model is made much wider.
- In very wide, shallow networks, the step-by-step method matched or beat the standard approach.
- Potential payoff: cheaper, less power-hungry AI training that's easier to spread across many computers.
Why It Matters
Could make future AI cheaper and less power-hungry, though wider models may need more memory.