Joshua Alfred

Data Scientist | MS Computer Science Student @ NYU | NYU Data Science Club

Jersey City, New Jersey, United States

About

👋🏼 Greetings! I am a second-year graduate student majoring in Computer Science at NYU, with a passion for Data Analytics, ML System Design, and Software Development. I'm a composed problem solver, with a zeal for researching and trying new methodologies to solve objectives efficiently. My research interests include DL architectures and LLM development. 🛠️ I am equipped with a collection of tech tools and domain knowledge in the realms of machine learning, data engineering, and cloud services. 📟 Languages: SQL, Python, Java, R, and JavaScript 🧰 Frameworks: Apache Spark, Airflow, Hadoop, PyTorch, PowerBI, React JS, Docker, and Jenkins ☁️ Cloud Services: AWS EC2, S3, RedShift, Lambda, and Elastic Beanstalk 💼 I have internship experience in managing ETL pipelines for real-time health monitoring and developing ML pipelines for credit risk assessment and insurance underwriting. Moreover, I have published research articles in novel DL techniques to solve automation in robotics and agriculture. I'm also a writer at the NYU Data Science Club, and I write blog articles about new topics, tech tips, and controversies within the data science industry. 🤝 I'm open to talks, feedback on my ongoing projects, and opportunities for internship or full-time. Let's connect!

Experience

  • Data Scientist at Mphasis
    Aug 2025 - Present · 1 yr

  • Writer | NYU Data Science Review at Data Science Club @ NYU
    Sep 2023 - May 2025 · 1 yr 9 mos

    • Author and Editor of blog posts and articles published in the NYU Data Science Review page in Medium.com • Event Coordinator and Volunteer for events conducted by the NYU Data Science Club

  • Research Assistant at New York University
    Aug 2024 - Jan 2025 · 6 mos

    Part-time Research Assistant at the System & Artificial Intelligence (SAI) Lab @ NYU, under Prof. Sai Qian Zhang. • Architectured a distributed self-speculative decoding approach to CodeLLaMA 7B model to improve inference on multiple edge devices • Utilized PyTorch’s Distributed Data Parallel framework to shard the LLM across edges, with a centralized server to compare draft and final tokens, rendering an improved inference time of 0.064s per token • Orchestrated batch jobs of LLM training and inference tests using NYU HPC clusters

  • Software Engineering Fellow at Headstarter
    Jul 2024 - Sep 2024 · 3 mos

  • Data Analyst | Mphasis Javelina at Mphasis
    Jun 2024 - Aug 2024 · 3 mos