Shouldn't this new reduction in thinking output make the model cheaper to operate? Could it be that the model is the same old LLM, with a bit newer architecture but still doing the same things, including huge amounts of thinking, just that its thinking is less clear to human readers?
Those tend to be under so-called ‘max’ modes, or ‘ultra’ - where you scale up inference time compute and then choose amongst your answers. Astra is new weights.
Comments
Shouldn't this new reduction in thinking output make the model cheaper to operate? Could it be that the model is the same old LLM, with a bit newer architecture but still doing the same things, including huge amounts of thinking, just that its thinking is less clear to human readers?
Those tend to be under so-called ‘max’ modes, or ‘ultra’ - where you scale up inference time compute and then choose amongst your answers. Astra is new weights.
Thinking output is already a relatively small part of the total input compared to raw code/text files.