top of page
Adiuvo Engineering & Training logo
MicroZed Chronicles icon

MicroZed Chronicles: Local AI, Skills and a Second Run

5 hours ago
6 min read

FPGA Horizons London- October 6th and 7th 2026 - get Tickets here.

The $99 Artix UltraScale+ Explorer Board - learn more here


Last week we looked at whether a language model running entirely on a machine on my desk could generate RTL that compiles. The short answer was yes, with the caveat that a compile pass means the code parses and elaborates and nothing more or as one comment on linkedin put it


“It compiles, so let's ship it!! Exactly my sense of humour”


That of course was only the beginning, does the code actually function, and does it survive lint, are important questions to determine if we have a really viable solution here.

To find out I took the final RTL from each model, wrote a self-checking test bench for every task, and ran the lot through QuestaSim and Blue Pearl.


The first answer was no, not really. The second answer, once I fixed the way I was asking, was a good deal better.


What the Models Were Asked For


The six tasks are deliberately ordinary. A parameterisable counter with an unsigned output. A two-flop CDC synchroniser with ASYNC_REG attached to the registers. An AXI4-Lite slave with four registers that honours the ready and valid handshakes on every channel. An 8N1 UART transmitter with a programmable divisor and a busy flag. A synchronous FIFO with wrap-bit pointers and a clocked block RAM read. And a bug fix on a small block with a broken reset and a registered mux.


The First Results


We have 84 designs, seven models by six tasks by two languages. Each task has one test bench used for every model in both languages, so everyone faces the same stimulus and checking. Simulation was QuestaSim 2025.3, a pass needed every check and every parameter variant to pass, each source was reviewed for what simulation cannot see, and every design went through Blue Pearl. Across the run the testbenches recorded over 65,000 assertion evaluations.


Of the 84 designs, 33 passed the bench. Twenty three compiled, ran and failed a test, 21 failed to compile in Questa, four failed to elaborate and three were missing ports the bench needed. Seven bench passes missed a stated requirement, leaving 26 designs, 31 percent, that did what was asked.


Model

VHDL bench

SV bench

Meets request / 12

Lint errors (VHDL / SV)

bluepearl-rtl

3 / 6

4 / 6

7

6 / 0

Qwen3.8-Flash-Next (125B MoE)

3 / 6

3 / 6

5

7 / 3

Gemma 4 12B

2 / 6

4 / 6

4

37 / 0

Qwen3-Coder 30B (MoE)

3 / 6

2 / 6

4

13 / 3

Qwen3.8 27B (dense)

2 / 6

3 / 6

3

20 / 0

Qwen 3.5 9B

1 / 6

1 / 6

2

39 / 75

Qwen 3.5 4B

2 / 6

0 / 6

1

82 / 78

Task

Bench passes / 14

Meets request / 14

counter

13

8

bugfix

11

11

cdc_sync

5

3

uart_tx

3

2

fifo

3

2

axi_lite

0

0

The counter and bug fix were essentially solved. Not one AXI4-Lite slave passed. The failures clustered into four patterns: handshakes, with BVALID asserted regardless of BREADY and read data changing under back pressure; FIFO reads implemented as a combinational lookup rather than the clocked RAM read requested; ASYNC_REG declared but not attached, or attached to one stage only; and UART bit timing a cycle out. Then the counter, which five of seven VHDL models declared as std_logic_vector when unsigned was asked for.


Last week's compile check was also too kind. GHDL and Icarus accepted seven designs Questa refused, four of them driving one variable from both always_comb and always_ff.


The Real Problem Was Me


Every failure on that list is a piece of knowledge a frontier model carries around in its head. What BVALID waits for. Where a VHDL attribute specification goes. Which edge a block RAM read lands on. A 4B or 12B model on a desktop does not have that context, and I had been prompting it as if it did.


So I wrote skills. One language skill each for VHDL and Verilog, covering the coding conventions, type rules, attribute syntax and inference patterns the compiler failures showed were missing. Then a protocol skill for each language, defining the behaviour of the AXI4-Lite handshakes, the UART framing and timing, the FIFO pointer and read structure and the synchroniser attribute rules. The skills give the model the information to steer by rather than leaving it to guess.


The Second Results


With the skills in place the whole run was repeated on the AMD Halo, same models, same tasks, same testbenches.

Model

VHDL bench

SV bench

Meets request / 12

Lint errors (VHDL / SV)

bluepearl-rtl

6 / 6

6 / 6

11

0 / 0

Qwen3.8-Flash-Next (125B MoE)

6 / 6

6 / 6

11

0 / 0

Qwen3.8 27B (dense)

6 / 6

6 / 6

11

0 / 0

Gemma 4 12B

5 / 6

5 / 6

9

0 / 0

Qwen3-Coder 30B (MoE)

4 / 6

6 / 6

9

0 / 0

Qwen 3.5 9B

5 / 6

3 / 6

7

0 / 0

Qwen 3.5 4B

4 / 6

3 / 6

6

25 / 0

Task

Bench passes / 14

Meets request / 14

counter

14

8

bugfix

11

10

cdc_sync

12

12

uart_tx

11

11

fifo

14

14

axi_lite

9

9

Bench passes went from 33 to 71 of 84, and designs that meet the request from 26 to 64, 31 percent to 76 percent. Three models, bluepearl-rtl, Flash-Next and the dense Qwen3.8 27B, now pass every bench in both languages. AXI4-Lite went from nothing to nine of fourteen, and the five that passed also cleared a supplementary 907-check suite of strobe masks, channel skew and response causality. The FIFO went to fourteen of fourteen with every read clocked, and every synchroniser now carries ASYNC_REG. Lint errors fell from 363 to 25, all of them in the 4B model's VHDL.


Not everything moved. Six of seven VHDL counters still declare the output as std_logic_vector rather than unsigned, so the request check keeps them at eight. Two synchronisers clock the chain from the source domain. The 4B model still cannot write an AXI slave that compiles.


The testbenches did not change, the models did not change, and the machine did not change. The only difference was telling the model what a competent engineer already knows.


What Next


The next step is to close the loop, feeding bench failures back to the model the way compiler errors were, and then synthesis, because a design that passes a bench and loses its synchroniser placement has not really passed.


The RTL, testbenches, skills, logs and lint reports for both runs are in the GitHub repository, with the scripts to rerun them with the licensed tools.


This is going to be a journey, but I think a worthwhile one.


FPGA Conference

FPGA Horizons London- October 6th and 7th 2026 - get Tickets here.


FPGA Journal

Read about cutting edge FPGA developments, in the FPGA Horizons Journal or contribute an article.


Workshops and Webinars:

If you enjoyed the blog why not take a look at the free webinars, workshops and training courses we have created over the years. Highlights include:



Boards

Get an Adiuvo development board:

  • Adiuvo Embedded System Development board - Embedded System Development Board

  • Adiuvo Embedded System Tile - Low Risk way to add a FPGA to your design.

  • SpaceWire CODEC - SpaceWire CODEC, digital download, AXIS Interfaces

  • SpaceWire RMAP Initiator - SpaceWire RMAP Initiator,  digital download, AXIS & AXI4 Interfaces

  • SpaceWire RMAP Target - SpaceWire Target, digital download, AXI4 and AXIS Interfaces

  • Other Adiuvo Boards & Projects.


Embedded System Book   

Do you want to know more about designing embedded systems from scratch? Check out our book on creating embedded systems. This book will walk you through all the stages of requirements, architecture, component selection, schematics, layout, and FPGA / software design. We designed and manufactured the board at the heart of the book! The schematics and layout are available in Altium here.  Learn more about the board (see previous blogs on Bring up, DDR validation, USB, Sensors) and view the schematics here.


All words in this blog were written by a human.

bottom of page