If they every gave really finegrained constraints you could constrain to subsets of tokens and extract the logits a lot cheaper than by random sampling limited to a few top choices and distill claude at a much deeper level. I wonder if that plays into some of the restrictions.
Comments
If they every gave really finegrained constraints you could constrain to subsets of tokens and extract the logits a lot cheaper than by random sampling limited to a few top choices and distill claude at a much deeper level. I wonder if that plays into some of the restrictions.
That makes sense, and if that's the reason it's another vote for open models