POLYTONE — the AI-native programming language

The generation benchmark

“Leading LLM programming language” is an empirical claim: a model receives a task, writes POLYTONE, and ptc test judges — deterministically, no human in the loop. Two numbers per task, both first-class: pass@1 (the first candidate passes) and pass@2e (after a failure the model sees the actual error and gets one repair attempt — POLYTONE’s thesis is that its errors teach).

Methodology

Results

claude-fable-52026-08-28 (run 10, toolchain 0.41.266 + card v16 evaluated, corpus 48 — first tier-L measurement) · workflow-harness

pass@1 43/48 · pass@2e 47/48 · tokens-to-green median 85907 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status85999
audio_probe171906
badge_page86100
bit_parity85754
clock_iso85832
config_port85828
content_tag85964
csv_report86790
csv_totals85849
dice_walk85948
env_banner85907
env_mode85730
fetch_all85982
grade_book85781
hex_dump85699
histogram85694
image_probe85776
inventory_load87759
json_pluck85791
log_scan85984
mesh_probe85746
model_parts86474
money_order86103
parse_point85767
pkg_check173085
poster_film86038
ranked_pick86253
rider_film86089
row_sums85703
run_length85984
save_report85804
season_label85747
set_overlap85861
shape_area85915
shape_report86039
stat_kit86021
stereo_field85865
swap_pairs85895
task_batch85904
top_scorer85998
total_of85797
tune_mix172459
uniques85846
video_probe172125
web_toc
week_later85782
wipe_reveal86027
word_stats85797
claude-fable-52026-08-28 (run 9, toolchain 0.40.259 + card v15 evaluated, corpus 36) · workflow-harness

pass@1 27/36 · pass@2e 35/36 · tokens-to-green median 82368 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status82378
audio_probe82740
bit_parity82299
clock_iso82204
config_port165699
content_tag165015
csv_totals82589
dice_walk82314
env_mode82216
grade_book82475
hex_dump82157
histogram82206
image_probe82319
json_pluck
log_scan82368
mesh_probe82239
money_order82484
parse_point165288
rider_film167221
row_sums82081
run_length82364
save_report164963
season_label164883
set_overlap82231
shape_area82313
stereo_field82556
swap_pairs82263
task_batch82274
top_scorer82351
total_of82401
uniques82213
video_probe82263
web_toc82543
week_later165795
wipe_reveal166712
word_stats82440
claude-fable-52026-08-23 (run 8, toolchain 0.40.252 + card v14 evaluated, corpus 36) · workflow-harness

pass@1 33/36 · pass@2e 36/36 · tokens-to-green median 68924 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status69726
audio_probe69347
bit_parity68601
clock_iso68554
config_port70896
content_tag68755
csv_totals69437
dice_walk68697
env_mode68547
grade_book68429
hex_dump68498
histogram68509
image_probe68873
json_pluck70577
log_scan70854
mesh_probe69024
money_order69621
parse_point138572
rider_film140779
row_sums68505
run_length68847
save_report69002
season_label68586
set_overlap68362
shape_area68591
stereo_field69319
swap_pairs68555
task_batch68958
top_scorer69066
total_of68887
uniques68600
video_probe68890
web_toc70251
week_later69937
wipe_reveal144018
word_stats68978
claude-fable-52026-08-14 (run 7, toolchain 0.35.239 + card v14, corpus 36) · workflow-harness

pass@1 34/36 · pass@2e 36/36 · tokens-to-green median 62827 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status63220
audio_probe63857
bit_parity62081
clock_iso62282
config_port65270
content_tag62631
csv_totals63061
dice_walk62228
env_mode62063
grade_book63908
hex_dump62190
histogram61990
image_probe62940
json_pluck64443
log_scan65030
mesh_probe63003
money_order63404
parse_point61959
rider_film65927
row_sums62012
run_length62627
save_report125537
season_label62479
set_overlap61873
shape_area62239
stereo_field63284
swap_pairs62142
task_batch62198
top_scorer63275
total_of62348
uniques62357
video_probe62295
web_toc64580
week_later64442
wipe_reveal127583
word_stats62827
claude-fable-52026-08-11 (run 6, toolchain 0.35.232 + card v14) · workflow-harness

pass@1 33/33 · pass@2e 33/33 · tokens-to-green median 56514 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status56619
audio_probe56530
bit_parity56499
clock_iso56460
config_port56582
content_tag56516
csv_totals56600
dice_walk56534
env_mode56451
grade_book56524
hex_dump56416
histogram56418
image_probe56520
json_pluck56554
log_scan56521
mesh_probe56491
money_order56589
parse_point56529
row_sums56432
run_length56514
save_report56541
season_label56524
set_overlap56443
shape_area56510
swap_pairs56474
task_batch56496
top_scorer56504
total_of56514
uniques56444
video_probe56458
web_toc56493
week_later56547
word_stats56519
claude-fable-52026-08-11 (run 5, toolchain 0.35.230 + card v13) · workflow-harness

pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 56471 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status56521
audio_probe56471
bit_parity56435
clock_iso56429
config_port56531
content_tag56456
csv_totals56583
dice_walk56480
env_mode56431
grade_book56488
hex_dump56404
histogram113469
image_probe56482
json_pluck114016
log_scan56471
mesh_probe56443
money_order56542
parse_point56485
row_sums56378
run_length56476
save_report56485
season_label56426
set_overlap56410
shape_area56477
swap_pairs56435
task_batch56467
top_scorer56445
total_of56485
uniques56389
video_probe56437
web_toc56436
week_later56487
word_stats113864
claude-fable-52026-08-11 (run 4, toolchain 0.35.226 + card v12) · workflow-harness

pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 56431 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status56462
audio_probe56446
bit_parity56384
clock_iso56381
config_port56513
content_tag56411
csv_totals56526
dice_walk56460
env_mode56359
grade_book56416
hex_dump56351
histogram113382
image_probe56428
json_pluck113895
log_scan56431
mesh_probe56405
money_order56503
parse_point113770
row_sums56342
run_length56432
save_report56457
season_label56378
set_overlap56434
shape_area56434
swap_pairs56392
task_batch56416
top_scorer56405
total_of56435
uniques56353
video_probe56387
web_toc56404
week_later56446
word_stats56439
claude-fable-52026-08-08 (run 3, toolchain 0.34.210 + card v12) · workflow-harness

pass@1 30/33 · pass@2e 32/33 · tokens-to-green median 50669 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status50695
audio_probe50681
bit_parity50645
clock_iso50633
config_port50746
content_tag50649
csv_totals50765
dice_walk50689
env_mode101841
grade_book50682
hex_dump50580
histogram101799
image_probe50668
json_pluck
log_scan50685
mesh_probe50636
money_order50746
parse_point50669
row_sums50607
run_length50662
save_report50678
season_label50629
set_overlap50598
shape_area50700
swap_pairs50618
task_batch50660
top_scorer50664
total_of50686
uniques50607
video_probe50624
web_toc50641
week_later50706
word_stats50684
claude-fable-52026-08-08 (run 2, toolchain 0.34.203 + card v10) · workflow-harness

pass@1 30/33 · pass@2e 33/33 · tokens-to-green median 50655 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status50687
audio_probe50666
bit_parity50635
clock_iso50616
config_port50727
content_tag50640
csv_totals50757
dice_walk50683
env_mode50599
grade_book50667
hex_dump50572
histogram101773
image_probe101976
json_pluck102320
log_scan50677
mesh_probe50632
money_order50731
parse_point50661
row_sums50601
run_length50655
save_report50677
season_label50638
set_overlap50598
shape_area50652
swap_pairs50626
task_batch50647
top_scorer50626
total_of50665
uniques50587
video_probe50625
web_toc50633
week_later50671
word_stats50671
claude-fable-52026-08-08 · workflow-harness

pass@1 21/33 · pass@2e 30/33 · tokens-to-green median 50372 (harness tokens: combined incl. the agent system prompt — comparable across harness runs, not to API usage)

taskpass@1pass@2etok→green
api_status50401
audio_probe50368
bit_parity50324
clock_iso50297
config_port101628
content_tag50335
csv_totals50450
dice_walk50376
env_mode50297
grade_book50356
hex_dump50281
histogram101160
image_probe101370
json_pluck50385
log_scan101442
mesh_probe
money_order50422
parse_point101544
row_sums101220
run_length101400
save_report50375
season_label50314
set_overlap50285
shape_area50351
swap_pairs
task_batch101356
top_scorer50334
total_of101495
uniques50282
video_probe50308
web_toc
week_later50369
word_stats50369

The first delta proof (replays)

The phase's first failure dataset is the corpus-authoring session itself (Sprint 133) — an LLM writing POLYTONE cold, its first attempts recorded verbatim. Sprint 138 replays them against the hardened toolchain, CI-gated (gen_replays.rs):

The tier-L delta (Sprints 267–268)

Run 10 (corpus 48) had five first-attempt failures; two of them are the first multi-module replays (pkg_check_attempt1/, tune_mix_attempt1/ — directory form, every file its own entry) and web_toc_attempt2_run10.pt the residual. Sprint 268 rewrote two teaching errors from them (a partial variant pattern shows the exact named form; an element method on a list names the way to an element) and measured COLD over the same five tasks: 3/5 pass@1 (0/5 in run 10). The residual named the last gap — there was no way to write an arm that does nothing — so pass became the empty statement; the same candidates re-judged under it read 4/5, the repair round 5/5. All three replays are pinned in gen_replays.rs.

Fresh pass@1/pass@2e runs against frontier models stay local and BYO-key — this section shows only what has actually run.

The corpus (48)

17 × tier S · 19 × tier M · 12 × tier L (multi-module: one file per module, each verified as its own entry) — each task: a prompt, a hidden judge, a reference solution.

tasktierarea
api_statusMCapabilities (mocked)
audio_probeMMedia codecs
badge_pageLModules (tier L)
bit_parityMBitwise & encoding
clock_isoSCapabilities (mocked)
config_portMCapabilities (mocked)
content_tagSBitwise & encoding
csv_reportLModules (tier L)
csv_totalsMData formats
dice_walkSCapabilities (mocked)
env_bannerLModules (tier L)
env_modeSCapabilities (mocked)
fetch_allLModules (tier L)
grade_bookSCollections
hex_dumpSBitwise & encoding
histogramSCollections
image_probeMMedia codecs
inventory_loadLModules (tier L)
json_pluckMData formats
log_scanSData formats
mesh_probeMMedia codecs
model_partsLModules (tier L)
money_orderMMethods & traits
parse_pointSErrors & Result
pkg_checkLModules (tier L)
poster_filmLModules (tier L)
ranked_pickLModules (tier L)
rider_filmMMedia codecs
row_sumsSCollections
run_lengthMPrelude & text
save_reportMCapabilities (mocked)
season_labelSMethods & traits
set_overlapSCollections
shape_areaSMethods & traits
shape_reportLModules (tier L)
stat_kitLModules (tier L)
stereo_fieldMMedia codecs
swap_pairsSGenerics
task_batchMAsync & Tasks
top_scorerSCollections
total_ofSErrors & Result
tune_mixLModules (tier L)
uniquesMGenerics
video_probeMMedia codecs
web_tocMMedia codecs
week_laterMData formats
wipe_revealMMedia codecs
word_statsSPrelude & text