Riya Basak

Riya Basak

I am a BSc (Hons) Computer Science with Artificial Intelligence graduate from the University of Hertfordshire, preparing for integrated MS–PhD research in artificial intelligence. My research goal is to develop mathematically grounded world models for general, reliable, and controllable intelligence: learning structured predictive representations of the physical world for reasoning, planning, and safe action. I am particularly interested in self-supervised predictive learning, compositional and causal world modeling, uncertainty-aware planning, and embodied AI.


News

  • Jul 2026–Present: R&D contributions in the Hugging Face open-source ecosystem, focused on world models and embodied AI, including LeRobot-related reproducibility, evaluation, benchmark compatibility, runtime correctness, and closed-loop evaluation. 🤗
  • Aug 2026: LAADAN-AC: Beyond Survival in Admissible Offline Treatment-Policy Learning accepted for oral presentation at ICATAS 2026! 🎉
  • Aug 2026: Awarded the University Graduation Prize by the Board of Examiners for the highest achievement at Level 6 in BSc (Hons) Computer Science (Artificial Intelligence), University of Hertfordshire. 🏆
  • Aug 2026: Released the preprint Mitigating Shortcut Learning in Brain Tumour MRI Classification on Zenodo. 📄
  • May 2026: Graduated from the University of Hertfordshire with First Class Honours in BSc (Hons) Computer Science with Artificial Intelligence. 🎓


Publications

Mitigating Shortcut Learning in Brain Tumour MRI Classification

Riya Basak
Preprint / Dissertation Research, 2026

Proposed Pathology-Focused Disentanglement (PFD) and Guided Semantic Token Evolution (GSTE) for shortcut-learning mitigation without segmentation masks in a ResNet50V2–RViT hybrid, with leakage-aware preprocessing, systematic ablation analysis, Grad-CAM++, attention rollout, and predictive-uncertainty analysis.

LAADAN-AC: Beyond Survival in Admissible Offline Treatment-Policy Learning

Riya Basak, Manal Helal
ICATAS 2026 Oral Presentation

Extended an admissibility-aware offline actor-critic framework with hard action masking, twin critics, conservative critic regularisation, expert-policy regularisation, smoothness shaping, Lagrangian cost control, component ablations, safety-failure analysis, and a cross-domain portability check.



Research Interests

  • World Models, Abstraction & Planning — structured latent state, transferable dynamics, causal abstraction, and long-horizon reasoning.
  • Self-Supervised Predictive Learning — latent predictive representations that capture controllable structure without reconstructing every observation detail.
  • Reliable & Controllable Intelligence — uncertainty-aware planning, explicit constraints, counterfactual evaluation, and safe action selection.
  • Embodied Intelligence — learning physical dynamics and transferable skills for agents that perceive, reason, plan, and act in changing environments.


Selected Research Projects

Admissibility-Aware Offline Actor-Critic Learning for Safer ICU-Sepsis Treatment Decisions

Safe Offline Deep Reinforcement Learning, 2026

Proposed LAADAN-AC, combining hard admissibility masking, twin critics, conservative critic regularisation, expert-policy regularisation, smoothness shaping, and Lagrangian cost control. Under the fixed benchmark protocol, the reported model achieved 0.7931 ± 0.0007 survival/return, 0.0000 inadmissibility, and 0.9525 expert-action match.

From World Models to Embodied Social Robots

Embodied AI · Social Robotics · World Models · Ongoing

Building a LeWM- and DreamerV3-based world-model research stack for agents that reason over predicted physical and social futures. The project proposes KCON (Kinetic Counterfactual Oversight Nexus) as a compact candidate executive layer for failure-aware action selection using counterfactual futures, uncertainty/novelty handling, and minimal intervention.

Same Rules, New Worlds

World Models · Causal Dynamics · Compositional Generalisation · Ongoing

Investigating why predictive models that perform well in-distribution can fail to transport learned dynamics across changes in context, appearance, embodiment, and long-horizon interaction. The current direction separates persistent state from transferable dynamics and evaluates compositional and interventional generalisation under controlled shifts and ablations.



Education

University of Hertfordshire — BSc (Hons) Computer Science with Artificial Intelligence, First Class Honours, 2026
Cumulative GPA: 4.38/4.50 · Level 6 GPA: 4.44/4.50
Ranked 1st in the Final Year Project across AI and all assessed streams · Project Report — 93% · Demo/Viva — 100%
Ranked 1st in the AI research-project module · LAADAN-AC Coursework — 91%



Honours & Awards

University Graduation Prize, University of Hertfordshire, 2026 🏆
Awarded by the Board of Examiners for the highest achievement at Level 6 in the BSc (Hons) Computer Science (Artificial Intelligence) programme.



Research Experience

  • Hugging Face Open-Source Ecosystem — R&D Contributor · Jul 2026–Present
    Contributing to open-source world models and embodied AI community work, including LeRobot-related reproducibility, evaluation, benchmark compatibility, runtime correctness, regression testing, and closed-loop evaluation.
  • Mercor — Artificial Intelligence Researcher, Mathematical Reasoning · Sep–Nov 2025
    Investigated failure modes in AI mathematical reasoning by formalising assumptions, proof obligations, case splits, edge cases, and recurrent proof errors into structured evaluation criteria.
  • Cubble — Full Stack AI Intern, ML Systems Research & Evaluation · Aug–Sep 2025
    Built model-review, safety-flag, logging, and failure-analysis tools connecting frontend inspection with backend AI services and reproducible testing.
  • Cubble — Software Engineer Intern, Machine Learning Research · Jul–Aug 2025
    Compared ResNet50- and CLIP-based approaches for safe/unsafe image classification and achieved 97.5% accuracy on unseen images with stable validation behaviour.


Research Direction & Impact

My long-term goal is to develop AI systems that can learn reusable models of how the world changes, separate causal mechanisms from nuisance variation, simulate plausible futures before acting, and remain controllable when predictions are uncertain or candidate actions are unsafe. The research objective is not benchmark performance alone, but reliable generalisation and decision-making when environments, tasks, or embodiments change.