Islamabad, Pakistan

Muhammad Uzair. Reinforcement learning researcher working on multi-agent DRL and human-guided policy optimization.

Two IEEE publications on cooperative multi-agent RL for wireless systems. Mitacs Globalink researcher at McMaster University (2025), advised by Dr. Istvan David, on integrating human advisory signals into PPO via subjective logic. Open to research-focused Masters opportunities, available to start any upcoming term.

2
IEEE publications
Mitacs
Globalink 2025
3.61
CGPA / 4.00
2025
NUST B.E. SE

01 / Research

Research interests.

I am looking for research-focused Masters opportunities in machine learning, available to start any upcoming term. My focus is reinforcement learning under realistic constraints: partial observability, multi-agent coordination, and learning from limited or noisy human feedback.

02 / Background

About.

I graduated from NUST in 2025 with a B.E. in software engineering and a 3.61 / 4.00 CGPA. My research track started in my third year at the Information Processing and Transmission Lab under Prof. Dr. Syed Ali Hassan, where I co-authored two IEEE papers on multi-agent DRL for wireless systems.

In summer 2025 I was a fully funded Mitacs Globalink research intern at McMaster University with Dr. Istvan David at the SSM Lab. I worked on guided policy optimization: integrating human advisory signals into PPO via subjective logic belief modeling on Gymnasium PacMan, with measurable convergence gains over baseline PPO.

Alongside research, I work as an AI engineer at Adept Tech Solutions on voice and email agents. This pays the bills and keeps me close to production LLM systems, which is informing my interest in RLHF and reward modeling research.

03 / Peer reviewed

Publications.

  1. [01]

    Multiagent Reinforcement Learning for Joint Spectrum and Energy Optimization in CR-NOMA Enabled Internet of Unmanned Agents

    Saleha Ahmed, Muhammad Uzair, Syed Asad Ullah, et al.

    IEEE Internet of Things Journal · 2025

    A cooperative multi agent DRL framework for CR-NOMA IoT, where distributed agents jointly learn spectrum access and power control policies under partial observability.

    pdf doi
  2. [02]

    Energy Efficient Uplink Communications for Wireless Powered Networks with EH Diversity: A DRL-driven Strategy

    Saleha Ahmed, Muhammad Uzair, Syed Asad Ullah, et al.

    IEEE International Conference on Communications (ICC) · 2025

    DRL driven transmit power control for energy harvesting uplink nodes, evaluated against MRC, SC, and EGC diversity combining schemes under Rayleigh fading.

04 / Selected work

Projects.

05 / Timeline

Experience.

  1. research

    Jun 2025 to Aug 2025

    Research Intern, Mitacs Globalink · McMaster University

    Hamilton, ON, Canada

    Advised by Dr. Istvan David · SSM Lab

    • Fully funded Mitacs Globalink internship on guided policy optimization in sequential decision making under partial observability.
    • Benchmarked REINFORCE and PPO on Gymnasium PacMan; tuned reward shaping and entropy regularization for stable convergence.
    • Designed a subjective-logic belief model that fuses human advisory signals with the policy gradient, achieving measurable convergence speedup over baseline PPO.
  2. research

    Jun 2024 to Sep 2025

    Research Collaborator · Information Processing and Transmission Lab, NUST

    Islamabad, PK

    Advised by Prof. Dr. Syed Ali Hassan

    • Co-authored two IEEE publications on multi-agent DRL for cognitive-radio NOMA and wireless powered networks.
    • Developed a cooperative MARL framework for joint spectrum access and power control in CR-NOMA IoT under partial observability.
    • Benchmarked DDPG, TD3, and PPO for continuous-action transmit power control under stochastic Rayleigh fading.
    • Analyzed MRC, SC, and EGC diversity combining schemes for energy harvesting uplink nodes.
  3. industry

    Nov 2025 to Present

    AI Engineer · Adept Tech Solutions

    Islamabad, PK

    • Engineering production LLM systems: voice agents on VAPI and Deepgram with sub-400ms transcription latency.
    • Multi-agent LLM orchestration over FastAPI microservices, plus a RAG retrieval layer on pgvector with 768-dimensional MPNet embeddings.
    • Operational context that keeps me close to alignment, reward modeling, and inference-time control as research questions.

06 / Academic

Education.

  1. education

    Nov 2021 to Jun 2025

    B.E. Software Engineering · National University of Sciences and Technology

    Islamabad, PK

    • CGPA 3.61 / 4.00 over 133 credit hours. Degree conferred 13 June 2025.
    • School of Electrical Engineering and Computer Science.
    • 4x FBISE HSSC merit scholarship recipient.

    Relevant coursework

    CS-368 Reinforcement Learning A
    CS-471 Machine Learning A
    MATH-361 Probability & Statistics A
    MATH-121 Linear Algebra & ODEs A
    MATH-352 Numerical Methods A
    MATH-232 Complex Variables & Transforms A
    CS-416 Large Language Models B+
    CS-250 Data Structures & Algorithms B+
    CS-251 Design & Analysis of Algorithms B+
    MATH-101 Calculus & Analytical Geometry B+

07 / Get in touch

Contact.

available

Open to research-focused Masters opportunities in machine learning, available to start any upcoming term (Winter, Spring, or Fall 2026 / 2027). My research focus is reinforcement learning, multi-agent systems, and learning from human feedback.

If you are a professor or graduate admissions reviewer, email is the fastest way to reach me and I respond within a day. My CV is linked below.