00_1 / Index Applied Scientist · Research Engineer Hamburg, DE First author · SECL · EMNLP 2026

Rissal Hedna

I work on LLM calibration, safety, and evaluation: making models honest about what they know. First author of SECL (EMNLP 2026), with production experience building and shipping large-scale generative and multi-agent systems.

Calibration Safety Evaluation LLMs
Open to research collaborations · Available for consulting
▶ Try the live SECL demo Get in touch ↗
↓ Scroll
LLM Calibration✦Test-Time Training✦Hallucination Detection✦Evaluation✦Uncertainty✦ LLM Calibration✦Test-Time Training✦Hallucination Detection✦Evaluation✦Uncertainty✦
0%
Fewer unsupported answers, customer-facing LLM features
0%
Lower end-to-end latency via eval and monitoring
0%
Fewer parameters, T5 trained from scratch
0×
Faster repeat renders via caching layer
0+
Active users on production generation
00_2

Latest

Aug 2026
SECL accepted at EMNLP 2026
Self-Calibrating Language Models via Test-Time Discriminative Distillation was accepted to EMNLP 2026 in Budapest (October 24–29). Read the paper ↗
New
Aug 2026
Darija LLM from scratch, notebooks open
Two Colab notebooks that build a Darija language model end to end: an 8K-token BPE tokenizer, then a 10.6M-parameter GPT trained on the Darija Wikipedia corpus. The recorded course goes up soon, the code is already public. Open the repo ↗
New
Aug 2026
Joined PHAROS Labs full time
Now a Forward Deployed AI Engineer at PHAROS Labs in Hamburg, working on agentic systems for medical writing and regulated life sciences.
New
2026
MSc defended, University of Hamburg
Finished the MSc in Intelligent Adaptive Systems with a final grade of 1.49, with the thesis work that became SECL.
2026
Invited talk at LLMDay Hamburg
Presented SECL and the generation–discrimination gap to the Hamburg LLM community. See the talk below
00_3

Focus

Language models are often confidently wrong, especially out of distribution. My work is about closing that gap: aligning what a model says it knows with what it actually does.

SECL is the current thrust. It calibrates a model at test time using its own P("True") signal on synthetic data, no labels and no human supervision. The same instinct runs through my production work: detect when a system is wrong, measure it, ship the fix.

Reliability diagram
MODEL CONFIDENCE → ACTUAL ACCURACY →
Perfect calibration Overconfident model
00_4

Selected Work

★ 01
SECL
Self-Calibrating Language Models via Test-Time Discriminative Distillation. A test-time training pipeline that calibrates an LLM's confidence from its own P("True") signal on synthetic data, with no labels or human supervision.
First authorAccepted · EMNLP 2026▶ Live demo
INTERACTIVE
Run it yourself
→
★ 02
Darija LLM
From Scratch
A two-part course for Arabic-speaking developers: code a BPE tokenizer from bytes up with an Arabic-aware split pattern, then build and train a 10.6M-parameter GPT on the Darija Wikipedia corpus. Includes a tokenizer ablation measuring bits-per-character, plus clean standalone implementations in under 550 lines. Adapted from Karpathy's minbpe and nanoGPT.
Python · PyTorchTokenizationTransformers2 Colab notebooksVideo course soon
COURSE
Run it in Colab
↗
03
Automatic AV Mixing
A Model for the Automatic Mixing of Multiple Video and Audio Clips. Published at the International Conference on Cyberworlds (CW).
First authorIEEE CW 2023
PAPER
2023
↗
04
Adversarial Patch Attacks
Simulated Adversarial Patch Attacks on Vision-Based Logistics Systems.
Co-authorVision · Robustness
PAPER
arXiv
↗
05
T5-Efficient
Trained T5 from scratch for git commit message generation, then cut parameters 81% (17.7M to 3.3M) via pruning and architecture work. Converged in under 12h on one GPU.
PyTorchTransformersBF16
PROJECT
↗
06
UniLLM
A dense-retrieval RAG assistant with async streaming and an automated crawler for German study and living data. Cut query resolution time 13% via better chunking and indexing.
QdrantLlamaIndexRAG
PROJECT
↗
University of Hamburg+PHAROS Labs+Fiindo+AdaLab+Telerobotics Lab+ University of Hamburg+PHAROS Labs+Fiindo+AdaLab+Telerobotics Lab+
00_5

Speaking

LLMDay Hamburg · 2026
Invited talk · LLMDay Hamburg

Self-Calibrating
Language Models

LLMs often assert falsehoods with full confidence, especially on out-of-distribution tasks. This talk walks through SECL: using the generation–discrimination gap to train a model to double-check itself at inference time.

With lightweight LoRA updates on late transformer layers, a model's verbalized confidence is aligned to its own P("True") signal, and entropy-based gating keeps the compute overhead low enough to deploy.

View certificate ↗
00_6

Experience

2026 → now
Forward Deployed AI Engineer
PHAROS Labs · Hamburg
Client-facing AI development on agentic systems for medical writing and regulated life sciences, turning research-grade methods into things that ship.
2025 — 2026
Graduate Researcher
University of Hamburg
Built SECL: test-time training that calibrates LLM confidence from the model's own P("True") on synthetic data, no labels needed. Beats verbalized-confidence baselines without their cost.
2024 — 2026
AI Engineer
PHAROS Labs · Hamburg
Shipped hallucination-detection layers for customer-facing LLM features, cutting unsupported answers 78%. Ran weekly customer sessions and built eval and monitoring that cut latency 55%.
2025
AI Engineer (Contract)
Fiindo · Hamburg
Built a full agentic pipeline solo, turning raw financial data into finished short-form videos (ingestion, narration, MANIM chart animation, render). A caching layer made repeat renders ~3x faster.
2023 — 2024
Machine Learning Developer
AdaLab UG · Hamburg
Raised peak-load throughput 60% on a multimodal generation platform and stabilized production video generation for 500+ active users.
2022
Visiting Research Assistant
Advanced Telerobotics Lab · USA
Built a C/Python CNN at sub-50ms inference on embedded NVIDIA Jetson for obstacle detection in GPS-denied environments, with LiDAR-vision sensor fusion under tight constraints.
00_7 / Contact
Let's
work
together
↗
© 2026 Mohamed Rissal HednaHamburg, DEP(True) ≈ calibrated