London Area, United Kingdom
TalkyBuddy (banibot.ru/talkybuddy) — Full-stack AI foreign language learning ecosystem with a focus on real-time voice interaction. - Engineered a cross-platform ecosystem (Web, iOS/Android, Telegram Mini App). Architected a modular core using Domain-Driven Design (DDD) to decouple complex linguistic logic from infrastructure. - Integrated WebRTC & Omni-modal AI: Designed a low-latency audio exchange system for real-time AI conversations, achieving sub-second response times (TTFT). - AI Agent Orchestration: Developed an adaptive assembly system for AI tutors with self-correction loops, leveraging iterative experience to improve lesson flow. Tech Stack: React Native, Next.js, Node.js (NestJS), WebRTC, GPT-Realtime-Mini, GPT-Realtime 1.5, GPT-Realtime 2, OpenAI/Gemini API.
- Built a multimodal voice-enabled multi-agent AI system for internal company workflows, including planning, project initiation, and facilitation. - Designed supervised agent delegation and Planner–Executor orchestration for complex task decomposition and execution. - Developed real-time voice interaction with Next.js, WebRTC, OpenAI Agents SDK, GPT-Realtime, GPT-4o/GPT-5.2, and Claude Opus-class models. - Implemented Graph-RAG / LightRAG retrieval, document tokenization, and automated embeddings generation for internal knowledge search. - Built a backend agent supervisor with hierarchical orchestration and real-time SSE synchronization. Engineered a Python/FastMCP server with multi-user Google OAuth, Gmail, and Google Calendar integrations. - Integrated MCP-based tools for email/document manipulation and Perplexity Sonar API for external search and deep research.
Development of an LLM-based projects - Worked on use cases in Prompt Engineering and researched generation capabilities using Anthropic. - Formed methods for "Chain of Thought" and prompt chaining to optimize model performance. - Delivered the "Text Analysis for violations", "Text Generation based on survey" features. - Delivered... (under NDA some projects) Tech stack: - Python, ReactJS/NextJS - Models/LLMs: Claude 3.5 Sonnet/Haiku, Llama 3.1-70b, Replicate.com models. - BERT/S-BERT/RoBERTa, T5, Embeddings generators - Prompt-engineering: Zero-Shot/Few-Shot Prompting, Chain of thought, Self-Reflection, Multi-modal LLM’s - Python, Pydantic, FastAPI, Grafana, Docker, K8S.
Focused on sound and image generation tasks, leveraging state-of-the-art techniques from arXiv and paperswithcode. Skilled in designing neural networks with PyTorch, preparing and visualizing data, and testing setups using Gradio. Expertise includes: Advanced architectures: Vision Transformer (ViT), U-Net, diffusion models, and Stable Diffusion (v3, XL). Sound/audio generation: VQ-VAE, hierarchical Transformer models (MusicLM, AudioLM, MusicGen, AudioGen). Model training: Backpropagation, gradient descent, loss tuning, and autoregressive modeling. Research outcomes include a project published on GitHub: https://github.com/applehawk/music-transformer-diffusion. Passionate about exploring cutting-edge generative models and their real-world applications.
Continuous delivery of features for B2BX, B2Core, and B2BinPay mobile apps for both iOS and Android platforms. Managed team 10+ members. All described apps available at the links below.