New training primitives
I added Muon, a method for updating a model as it learns, and ReLU², a small mathematical function used inside neural networks. Both are credited in MLX’s contributor acknowledgements.
See MLX acknowledgements
Open-source contributor · Apple MLX
Systems Engineer at Computacenter, specialising in machine learning engineering and MLOps. Machine Learning Research Engineer in my spare time.
I build machine-learning systems, software that learns patterns from examples—and help take them from experiments into reliable services. That work is called MLOps. In my free time, I contribute to open source and research ways to run and train AI on Apple chips. I’m a key contributor to Apple’s MLX toolkit and MLX-LM, which make it easier to run and train language models locally. I also maintain the training engine for models that work with images and text, and develop my own J.O.S.I.E. AI model family.01 / About
At Computacenter, I work as a System Engineer focused on machine-learning engineering and MLOps—the work of deploying, monitoring and maintaining software that uses AI. Outside work, I contribute to open source and conduct ML research. Apple’s MLX tools let developers build, train and run models on Apple chips. My contributions help those tools support more models and make it easier to teach them new tasks, including models that understand both images and words.
02 / Upstream
MLX is Apple’s open-source toolkit for machine learning. The tools around it help people run and train AI models directly on Apple chips. Here are some of my credited contributions.
I added Muon, a method for updating a model as it learns, and ReLU², a small mathematical function used inside neural networks. Both are credited in MLX’s contributor acknowledgements.
See MLX acknowledgementsMLX-LM is Apple’s toolkit for running and training text-generating AI models. I helped it support more than 20 model designs, and added ways to train every learned value in a model, choose between training methods, and track experiments with Weights & Biases, a tool for recording and comparing training runs. A model design is its internal blueprint; the list below links to the original model files and code changes.
Read MLX-LM creditsMLX-VLM works with models that handle both images and words. I’m the main maintainer of its training engine—the code that runs the learning process. I rebuilt that engine and added ORPO, a way to teach models from examples of preferred and less-preferred answers.
Explore MLX-VLMMLX Examples offers practical, runnable demonstrations. Its acknowledgements credit my work bringing several models to the examples and adding full training, which adjusts all of a model’s learned values.
Read MLX Examples creditsMLX-LM credits my work on these 25+ model variants from 15+ organizations. An architecture is a model’s internal design; weights are the learned values that let it work. “Weights” links to original model files. “PR” means pull request: a proposed code change and its review.
MiniCPM · MiniCPM3
Weights MiniCPM family · PRs #685 · #24
Qwen3Next
Weights Qwen3 Next · PRs #441 · #453
Helium
Weights Helium-1 preview · 2B · PR #1208
DeepSeek v4.1
Weights DeepSeek v4.1 Flash · PR #1895
dots.llm1
Weights dots.llm1.inst · PR #211
ERNIE 4.5 MoE
Weights ERNIE-4.5-21B-A3B · PR #267
Bailing MoE (Ling) · Bailing MoE Linear (Ling-Linear)
Weights Ling-lite · Ring Mini Linear · PRs #369 · #513
Klear
Weights Klear-46B-A2.5B · PR #437
Jamba
Weights Jamba-v0.1 · PR #544
Granite MoE
Weights Granite 4.0 H Small · PR #413
Mistral4
Weights Mistral Small 4 · PR #1012
LongCat
Weights LongCat Flash Chat · PR #423
Nemotron H
Weights Nemotron H 8B · PR #407
Apertus
Weights Apertus 8B · PR #421
Lille130m
Weights Lille 130M · PR #429
Qwen3Next · Qwen3 · Qwen3MoE
Weights Qwen3-Next · Qwen3 · PRs #441 · #42* · #199
TeleChat3
Weights TeleChat3 36B · PR #773
Mamba v3 is still in review — weights · open PR #1021.
Qwen3 note: PR #42 closed before it was merged. The project credits my help on Qwen3 and its version built from specialist submodels; PR #199 is a later accepted update to that version.
Read the MLX-LM acknowledgements03 / Research
A model with several specialist parts that can activate different experts for different inputs. DynaMoE studies whether choosing those specialists dynamically can help a model use its computing resources more efficiently.
Read the paperA method for changing selected model behaviors by editing its learned values, while checking that other abilities still work well.
Read the paperDirectional and Similarity-aware Latent Alignment (DSLA) studies how feedback about better answers can also shape the patterns a model forms internally while processing text.
Read the paper04 / Model family
J.O.S.I.E. is my family of downloadable AI models, exploring reasoning, honesty and independent judgment. J.O.S.I.E.-2 includes models with 2, 4 and 9 billion learned values, trained on Apple chips. Smaller versions, with fewer bits used to store each learned value, need less memory and are easier to run on a personal computer.
Read the J.O.S.I.E.-2 research note Visit the J.O.S.I.E. project page05 / Selected work
Tools for adapting language models to new tasks on Apple chips. Fine-tuning continues a model’s training with new examples; LoRA does this by learning a small add-on instead of changing every value in the original model.
Explore the projectTraining tools for AI models that work with both images and text. I maintain the code that runs their training and adaptation workflows.
Explore the projectAn open-source Mac app for adapting language models to new tasks, checking their results and managing models on Apple chips.
Explore the projectA personal model family exploring small language models, reasoning and assistant behavior, with releases and research notes on Hugging Face.
Explore J.O.S.I.E.A research tool for looking inside language models: how they represent words, which parts of the input they focus on, and how information moves through them.
Explore the projectA notebook that uses your PDFs as the source for summaries and creates audio versions with language and speech tools you choose.
Explore the projectResearch and tools for finding model edits that change selected behaviors while aiming to keep other abilities intact.
Explore the projectAn implementation of Kolmogorov–Arnold Networks, an alternative model design built from simple, learnable functions, using Apple’s MLX toolkit.
Explore the project06 / Career
System Engineer focused on building machine-learning software and keeping it reliable after deployment. MLOps covers the tools and practices for shipping, monitoring and maintaining AI systems.
In my free time, I contribute to open source and pursue ML research and engineering.
07 / References
08 / Elsewhere