Post by Shearman Chua

Senior Engineer, Digital Hub(AIDA-C3)

šŸš€ Introducing A2A Multimodal DeepAgent Chat I'm excited to share a project I've been working on — a full-stack multimodal AI agent system that brings together several cutting-edge technologies: šŸ¤– What it does: Chat with AI agents using text, images, and videos Agents can detect and classify objects in images/videos using YOLO and Vision Language Models Built on Google's A2A (Agent-to-Agent) protocol for standardized agent communication Supports streaming responses for real-time interaction šŸ› ļø Tech Stack: Frontend: React + Vite + Tailwind CSS Agent: Python + LangGraph DeepAgent + LangChain Tools: FastMCP (Model Context Protocol) for extensible tool integration Storage: MinIO for media handling with pre-signed URLs Observability: Arize Phoenix for tracing and debugging ✨ Key Features: šŸ“ø Upload images/videos and get AI-powered analysis šŸŽÆ Target detection & classification with threat assessment šŸ” Web search integration via DuckDuckGo šŸ“” Real-time streaming via Server-Sent Events 🐳 Fully containerized with Docker Compose Why A2A Protocol? The A2A protocol enables standardized communication between AI agents, making it easier to build interoperable multi-agent systems. This project demonstrates how to build a production-ready A2A agent with multimodal capabilities. Check out the repo: https://lnkd.in/grgQx-dH Would love to hear your thoughts and feedback! šŸ’¬ #AI #MachineLearning #LLM #MultimodalAI #AgentAI #OpenSource #Python #React #Docker #Langgraph #A2A #arizephoenix

Post content

Video Content