LocateAnything Explained: Parallel Box Decoding and how it makes visual grounding faster and more precise
Andrey Lukyanenko
LocateAnything Explained: Parallel Box Decoding and how it makes visual grounding faster and more precise Paper Project Demo Modern detection-and-grounding VLMs treat a bounding box as text: each box becomes a short string of coordinate tokens, decoded one at a time, left to right. This...