Towards Pixel-level VLM Perception via Simple Points Prediction
Tianhui Song
sthui
AI & ML interests
None yet
Recent Activity
upvoted a paper 1 day ago
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs authored a paper 2 days ago
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers upvoted a paper about 1 month ago
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers