Islamabad, Pakistan

Muhammad Uzair. Machine learning researcher working on LLM personalization and alignment, built on a reinforcement learning foundation.

I work on personalizing language models to individual users, at the memory level and at the weights level, and on how you evaluate a system that is supposed to keep changing. Currently building persistent per-user AI twins in production. Two IEEE publications on multi-agent deep RL for a hard constrained optimization problem, and a Mitacs Globalink term at McMaster University (2025) on integrating human advisory signals into PPO via subjective logic. Looking for a thesis-based MSc for 2027.

2
IEEE publications
Mitacs
Globalink 2025
3.61
CGPA / 4.00
2025
NUST B.E. SE

01 / Research

Research interests.

I am looking for a thesis-based MSc for 2027. My focus is LLM personalization: how a model adapts to one specific person over time, whether that adaptation belongs in memory or in weights, and how you evaluate it honestly once the target keeps moving. Reinforcement learning under realistic constraints, partial observability and noisy human feedback, is where I came from and still the toolkit I reach for.

02 / Background

About.

I graduated from NUST in 2025 with a B.E. in software engineering and a 3.61 / 4.00 CGPA. My research track started in my third year at the Information Processing and Transmission Lab under Prof. Dr. Syed Ali Hassan, where I co-authored two IEEE papers on multi-agent deep RL for a hard constrained optimization problem.

In summer 2025 I was a fully funded Mitacs Globalink research intern at McMaster University with Dr. Istvan David at the SSM Lab. I worked on guided policy optimization: integrating human advisory signals into PPO via subjective logic belief modeling on Gymnasium PacMan, with measurable convergence gains over baseline PPO.

Since then my centre of gravity has moved to LLM personalization. I now work on persistent per-user AI twins at Alexein AI : assistants that carry a durable model of one person across months, and that have to be evaluated and improved without a fixed target to score against. Before that I spent a year at Adept Tech Solutions building production voice and email agents, which is where the retrieval, orchestration, and evaluation problems stopped being abstract.

I am pursuing the same question on two tracks. At the memory level: extracting durable preferences from raw interaction history and keeping them honest as they drift. At the weights level: multi-LoRA composition and hypernetworks that generate a per-user adapter rather than fine-tuning one model per person.

03 / Peer reviewed

Publications.

  1. [01]

    Multiagent Reinforcement Learning for Joint Spectrum and Energy Optimization in CR-NOMA Enabled Internet of Unmanned Agents

    Saleha Ahmed, Muhammad Uzair, Syed Asad Ullah, et al.

    IEEE Internet of Things Journal · 2025

    A cooperative multi agent DRL framework for CR-NOMA IoT, where distributed agents jointly learn spectrum access and power control policies under partial observability.

    pdf doi
  2. [02]

    Energy Efficient Uplink Communications for Wireless Powered Networks with EH Diversity: A DRL-driven Strategy

    Saleha Ahmed, Muhammad Uzair, Syed Asad Ullah, et al.

    IEEE International Conference on Communications (ICC) · 2025

    DRL driven transmit power control for energy harvesting uplink nodes, evaluated against MRC, SC, and EGC diversity combining schemes under Rayleigh fading.

04 / Selected work

Projects.

05 / Timeline

Experience.

  1. research

    Jun 2025 to Aug 2025

    Research Intern, Mitacs Globalink · McMaster University

    Hamilton, ON, Canada

    Advised by Dr. Istvan David · SSM Lab

    • Fully funded Mitacs Globalink internship on guided policy optimization in sequential decision making under partial observability.
    • Benchmarked REINFORCE and PPO on Gymnasium PacMan; tuned reward shaping and entropy regularization for stable convergence.
    • Designed a subjective-logic belief model that fuses human advisory signals with the policy gradient, achieving measurable convergence speedup over baseline PPO.
  2. research

    Jun 2024 to Sep 2025

    Research Collaborator · Information Processing and Transmission Lab, NUST

    Islamabad, PK

    Advised by Prof. Dr. Syed Ali Hassan

    • Co-authored two IEEE publications on multi-agent DRL for cognitive-radio NOMA and wireless powered networks.
    • Developed a cooperative MARL framework for joint spectrum access and power control in CR-NOMA IoT under partial observability.
    • Benchmarked DDPG, TD3, and PPO for continuous-action transmit power control under stochastic Rayleigh fading.
    • Analyzed MRC, SC, and EGC diversity combining schemes for energy harvesting uplink nodes.
  3. industry

    Aug 2026 to Present

    AI Research Engineer · Alexein AI

    Islamabad, PK

    • LLM personalization in production: persistent per-employee AI twins that adapt to an individual user over months rather than a single session.
    • Agent evaluation and evolution. Rubric-based and LLM-as-judge evaluation with explicit attention to judge bias, feeding a staged evolution path from memory and retrieval through prompt adaptation to preference optimization.
    • Direct production counterpart to my research interest in personalization at both the memory level and the weights level.
  4. industry

    Nov 2025 to Aug 2026

    AI Engineer · Adept Tech Solutions

    Islamabad, PK

    • Engineered production LLM systems: voice agents on VAPI and Deepgram with sub-400ms transcription latency.
    • Multi-agent LLM orchestration over FastAPI microservices, plus a RAG retrieval layer on pgvector with 768-dimensional MPNet embeddings.
    • Operational context that kept me close to alignment, reward modeling, and inference-time control as research questions.

06 / Academic

Education.

  1. education

    Nov 2021 to Jun 2025

    B.E. Software Engineering · National University of Sciences and Technology

    Islamabad, PK

    • CGPA 3.61 / 4.00 over 133 credit hours. Degree conferred 13 June 2025.
    • School of Electrical Engineering and Computer Science.
    • 4x FBISE HSSC merit scholarship recipient.

    Relevant coursework

    CS-368 Reinforcement Learning A
    CS-471 Machine Learning A
    MATH-361 Probability & Statistics A
    MATH-121 Linear Algebra & ODEs A
    MATH-352 Numerical Methods A
    MATH-232 Complex Variables & Transforms A
    CS-416 Large Language Models B+
    CS-250 Data Structures & Algorithms B+
    CS-251 Design & Analysis of Algorithms B+
    MATH-101 Calculus & Analytical Geometry B+

07 / Get in touch

Contact.

available

Looking for a thesis-based MSc in machine learning for 2027. My research focus is LLM personalization: adapting models to individual users at the memory and weights level, and evaluating agents that are meant to keep changing. Reinforcement learning and learning from human feedback are the foundation underneath it.

If you are a professor or graduate admissions reviewer, email is the fastest way to reach me and I respond within a day. My CV is linked below.