Share your thoughts, 1 month free Claude Pro on us
See more
Home
/
Benchmarks
Text-to-score generation on 238 evaluation prompts
Loading...
3.48
Prompt Adherence
Text2Score
1.5976
2.0863
2.575
3.0637
May 13, 2026
Prompt Adherence
Readability
Musicality
Authenticity
Usability
Updated 2mo ago
Evaluation Results
Method
Method
Links
Prompt Adherence
Readability
Musicality
Authenticity
Usability
Text2Score
2026.05
3.48
3.98
3.52
3.13
3.44
ComposerX
LLM engine=GPT-5.1
2026.05
2.94
2.92
2.92
2.44
2.65
Midi-LLM
2026.05
1.67
1.79
1.69
1.44
1.52
Feedback
Search any
task
Search any
task