MOA Workshop · November 1, 2026 · Athens, Greece
NPU Model Optimization Competition
Optimize the inference of Gemma 4 12B, a multimodal LLM, on FuriosaAI's RNGD.
Overview
Participants start from a baseline implementation of Gemma 4 12B written for RNGD with the furiosa-opt toolchain and improve its inference performance. Participants do not need their own RNGD hardware: submitted programs are run on FuriosaAI's RNGD servers through FuriosaAI Arena.
The competition is held in two rounds. The Kernel Optimization Round targets three decoder-layer kernels of the model — the QKV projection, the attention output projection and the feed-forward block. Participants with the highest results advance to the Model Optimization Round, in which the full model is optimized end to end. The top three teams will have the opportunity to present their work at the MOA Workshop in Athens.
In both rounds, submissions must pass a correctness check and are then ranked by performance. Detailed rules and evaluation criteria are specified in the baseline repository.
Getting started
Tutorial session
MOA 2026 NPU Kernel Optimization Tutorial will be held online on Tuesday, September 15, 23:00–24:00 UTC (Wednesday, September 16, 08:00–09:00 KST). A recording will be shared afterwards for those who cannot attend live.
-
Register
Fill out the registration form by September 15, 2026.
-
Set up the baseline
Clone the baseline repository, a working implementation of Gemma 4 12B for RNGD. Its README is the competition guide: it covers the toolchain setup, the three target kernels, the permitted changes and the grading test. OPTIMIZATION.md walks through the optimization workflow. The furiosa-opt documentation is the companion programming guide, covering everything from the vISA programming model to scheduling and tuning.
-
Optimize and test
Improve the target kernels and check your work on real hardware as you go: the
rngdcommand-line client runs your code on FuriosaAI Arena, the RNGD evaluation server. Arena access is provided to registered participants. -
Submit
When you are ready, submit your entry as described in the baseline repository's README. Submissions that pass the correctness check are ranked by performance on the leaderboard, and top performers in the Kernel Optimization Round advance to the Model Optimization Round.
Technical questions are best asked on the FuriosaAI forums, where answers are shared with all participants.
Important dates
| Item | Date |
|---|---|
| Registration Period | Sep 1 – Sep 15, 2026 |
| Kernel Optimization Round | Sep 1 – Sep 25, 2026 |
| Tutorial Session (online) | Sep 15, 2026, 23:00 UTC |
| Finalists Announcement | Sep 30, 2026 |
| Model Optimization Round | Oct 1 – Oct 25, 2026 |
| Award Ceremony (time and venue TBA) | Nov 1, 2026 |
The award ceremony takes place at the MOA Workshop, held on November 1, 2026 at MICRO 2026 in Athens, Greece.
Prizes
- 1st place€5,000
- 2nd place€3,000
- 3rd place€2,000
Winners must attend the on-site award ceremony in Athens.
Resources
- Baseline repository Gemma 4 12B baseline for RNGD. The README also serves as the competition guide. It specifies the target kernels, permitted changes and correctness tolerances, and explains how to set up the toolchain, optimize a kernel and run the grading test. github.com/micro2026-moa/furiosa-opt-gemma4-12B
- furiosa-opt documentation Programming guide for the furiosa-opt toolchain in which the baseline is written. Covers setup, the vISA programming model, scheduling, and development tools such as the schedule viewer. developer.furiosa.ai/furiosa-opt/book
- FuriosaAI Arena Job scheduler for FuriosaAI's shared RNGD servers, on which submitted programs are executed. Jobs are submitted with the rngd command-line client; setup instructions are in the baseline repository. Access is provided to registered participants. arena.furiosa.ai