Can we be sure that research became part of their mainline model development process as opposed to being an interesting side-quest?
Are Gemini and DeepSeek and Llama and other strong coding models using the same ideas?
Llama and DeepSeek are at least slightly more open about their training processes so there might be clues in their papers (that's a lot of stuff to crunch through though).
Comments
Can we be sure that research became part of their mainline model development process as opposed to being an interesting side-quest?
Are Gemini and DeepSeek and Llama and other strong coding models using the same ideas?
Llama and DeepSeek are at least slightly more open about their training processes so there might be clues in their papers (that's a lot of stuff to crunch through though).