Return to Article Details NEXT-GPT: Any-to-Any Multimodal LLM for Unified Text, Image, Audio, and Video Understanding Download Download PDF