Menlo Park, California, United States
I'm proud of my R&D background in AI Agents, Machine Learning, Search, Computer Vision, and Data Mining. I started my career devoting myself in image understanding with web-scale of data at Microsoft Research. Together with my colleagues while as a core person, I proposed the idea of "annotation by search" and hands-on built a real-time image annotation system on top of 4 billion images and 10 machines. I'm currently working in the Bard team at Google DeepMind to explore capabilities of AI agents. We happily announced https://blog.google/products/gemini/google-gemini-deep-research/. In my early role as a researcher, I published 40+ papers on image understanding, image retrieval, and machine learning in top conferences and journals like CVPR, MM, WWW, T-PAMI etc.. I also authored a book (in Chinese) "Reviews of New Technologies for Internet Communications and Networking", together with my PhD advisor at Tsinghua University. So far I hold 17 granted US patents.
Agentic data curation for Physical AI
Taxonomy-based produnct understanding, taxonomy evolution
AI Agentic Gemini features
ViT-based solutions for digitizing the electric grid
* I worked in the image search ranking team to tackle the diversity problems. With improved diversity of search results, we can save users time in seeking for their interested information. * Search "jaguar" or "tesla model x" in images.google.com. The carousels on top of the result pages illustrate the outcome of some of our efforts.