It's definitely a work in progress - but something that active development is being focused around.
The way this is being handled in an upcoming update involves a few things - an OCR tool identifies math formulas, applies a bounding box and takes an image. That image gets sent to a multimodal-LLM which attempts to "describe" the formula reasonably.
While not yet perfect, this is something I anticipate to improve quite a bit soon.
The same approach is going to be applied to tables, graphs, figures, and images.
Comments
It's definitely a work in progress - but something that active development is being focused around. The way this is being handled in an upcoming update involves a few things - an OCR tool identifies math formulas, applies a bounding box and takes an image. That image gets sent to a multimodal-LLM which attempts to "describe" the formula reasonably. While not yet perfect, this is something I anticipate to improve quite a bit soon. The same approach is going to be applied to tables, graphs, figures, and images.