中文 English

Is Peeking at Your Opponent's Chessboard Immoral? I Used This Legendary Troll Question to IQ-Test 13 LLMs

Published: 2026-08-14 · 阅读量 --
AI LLM Ruozhiba Benchmark Fun Chinese Chess

Short version

I picked a classic trick question from Ruozhiba—the legendary Chinese forum famous for absurd brain-teasers—and asked it to 13 LLMs through my self-hosted NewAPI gateway: “While playing chess, is it immoral to peek at your opponent’s board?” The results were dramatic: 10 models spotted the trap (in chess, both players share ONE public board—there is no “opponent’s board” to peek at), 1 model failed at first then talked itself out of the hole, 2 models fell straight in and started moralizing, and 1 model simply told me to “refresh the page and try again.” This is an unofficial, just-for-fun IQ report. Every answer is a real API response, archived with screenshots.

Hero image of the Ruozhiba exam

Figure 1: A peeking eye stares at a chessboard that both players were already allowed to look at. Original illustration.

1. Background: Ruozhiba, the unofficial “final exam” for Chinese LLMs

A quick introduction for readers outside the Chinese internet. Ruozhiba (literally “the forum of silly questions”) is a Baidu Tieba community where users post questions that sound reasonable for half a second and then collapse into beautiful absurdity—things like “if I put my left shoe on my right foot, is my left foot now inside the right shoe?” The questions look like jokes, but they hide real logic traps, puns, and broken premises.

This was pure internet fun until 2024, when a team of researchers published a serious paper: COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning. They collected training corpora from different Chinese internet platforms and discovered something hilarious: models fine-tuned on Ruozhiba data achieved the highest scores in 8 of the evaluation tasks, beating corpora from encyclopedias, Q&A sites, and carefully curated datasets. The authors suspect Ruozhiba questions strengthen logical reasoning.

Screenshot of the COIG-CQIA paper on arXiv

Figure 2: Real screenshot of the COIG-CQIA arXiv page. Ruozhiba entering academic literature was one of the most surreal moments of the Chinese internet in 2024.

Screenshot of the COIG-CQIA dataset on Hugging Face

Figure 3: The dataset is public on Hugging Face with over 7,000 downloads last month. Real screenshot.

Since then, Ruozhiba questions have become a folk benchmark for Chinese LLMs: can a model survive a Ruozhiba question without embarrassing itself? That became the inspiration for this test.

2. The question: where exactly is the trap?

The exam question: “While playing chess, is it immoral to peek at your opponent’s board?”

The genius of this question is that its premise is broken. Chess is not poker. In poker, your opponent’s hand is hidden information, and peeking is absolutely cheating. But in Chinese chess and international chess, both players look at the same single board—every piece’s position is fully public from start to finish. There is no “my board” versus “the opponent’s board.” The pieces sit right there; how would you even play without looking?

An analogy any elementary schooler can follow: imagine the teacher writes the exam questions on the classroom blackboard, and both you and your deskmate look up at it. Someone asks: “Isn’t it immoral to peek at the blackboard your deskmate is looking at?” The blackboard belongs to the whole class. Looking at it isn’t peeking—it’s attending class.

So the correct way to answer is: first dismantle the non-existent premise of “peeking at the opponent’s board,” then explain what actually IS immoral (reading the opponent’s phone, notes, engine analysis, or prepared opening files—or peeking in hidden-information games like Kriegspiel, blindfold chess, or Stratego-style variants). Any model that immediately starts a morality lecture without questioning the premise has fallen into the pit.

Trap analysis diagram

Figure 4: The trap explained—chess is a perfect-information game played on one shared board. Original diagram.

3. Method: one question, 13 models, one gateway

The setup is simple. My self-hosted NewAPI gateway aggregates official models from multiple vendors. I enrolled 13 models in the exam: gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, deepseek-v4, deepseek-v4-pro, deepseek-v4-flash, glm-5.2, GLM-5V-Turbo, k3, k3-256k, kimi-k2p6, qwen3.8-max, and MiniMax-M3.

One roster note: the gateway also hosts a batch of latest-suffixed models (gpt-latest, glm-latest, kimi-latest, qwen-latest, minimax-latest, etc.) plus some custom DIY aliases I defined myself. These are just “stage names” pointing at some underlying model—testing them would be grading photocopies of the same paper, and an alias that silently retargets to a new version makes results impossible to reproduce. They were all disqualified; only models with concrete version names were allowed on stage.

Exam rules: the same question, the same temperature=0.7, called one by one through the same OpenAI-compatible endpoint. I recorded each model’s full answer, latency, and token usage, then graded manually. The core grading criterion was a single line: did the model spot the “one shared public board” trap?

Test method diagram

Figure 5: The pipeline—one question fanned out to 13 models via a self-hosted NewAPI gateway. Original diagram.

4. Report card: who actually thinks, who recites the textbook

The scoreboard first. Once again: entertainment-only scores, not an official evaluation.

Entertainment IQ leaderboard

Figure 6: The entertainment IQ leaderboard of this Ruozhiba exam. Original chart; the criterion is trap detection.

4.1 Honor roll: dismantling the trap in one sentence

deepseek-v4-flash (98) submitted my favorite answer. It opened with “It depends,” followed immediately by: “In a normal open game, the chessboard is shared by both players and fully public—you’re supposed to look at it. Looking at the board to plan your move isn’t peeking; it’s called playing chess.” It closed with a punchline: “Open games need no peeking; peeking in hidden games is cheating.” Clean, correct, and witty.

Real answer from deepseek-v4-flash

Figure 7: Real API response from deepseek-v4-flash—the fastest trap detection and the most stylish answer.

k3 (97) was the most charming paper of the day. It opened with a laugh—“Ha, fun question—but when playing chess, there’s actually no such thing as ‘peeking at the opponent’s board’!"—then listed three reasons why looking at the board is mandatory, four genuinely违规 behaviors (engine assistance, peeking at notes, taking hints from bystanders, swapping pieces), and finished by asking back: “Did you run into a specific situation, or are you just joking with friends?” Trap dismantled, lesson delivered, conversation kept alive.

Real answer from k3

Figure 8: Real response from k3—dismantling the trap with a sense of humor.

deepseek-v4-pro (96) also nailed it in one sentence: “In standard Chinese or international chess, the board is public; both players look at the same board… so there is no such thing as peeking at the opponent’s board.” It then enumerated the genuinely immoral cases with clear structure.

Real answer from deepseek-v4-pro

Figure 9: Real response from deepseek-v4-pro—breaking the premise first, then discussing cases.

The gpt-5.6 trio (sol 96 / luna 95 / terra 94) all passed with a strikingly consistent style: fix the premise first, list the real cheating scenarios, close with one compact summary. sol’s closer—“Information the rules allow you to see, look freely; information the rules require to stay hidden, peeking is cheating”—was the crispest of the day, delivered with an emoji. terra offered the most human advice: “If you’re just curious in a casual game, simply ask: may I look?” Same teacher, same answer template—and reassuringly stable.

Real answer from gpt-5.6-sol

Figure 10: Real response from gpt-5.6-sol—compact, correct, and emoji-friendly.

k3-256k (96) and kimi-k2p6 (95) were equally decisive. k3-256k had the most vivid analogy: “What you’d see by ‘peeking’ at the opponent’s board is exactly what you see by looking openly.” kimi-k2p6 used three neat subheadings to explain the truly hateful behaviors (hidden chess, blindfold chess, snitching spectators), closing with an elder-statesman line: “Losing a game is fine; losing your integrity is not—a win by peeking is more shameful than a loss.”

Real answer from kimi-k2p6

Figure 11: Real response from kimi-k2p6 (excerpt)—well structured, with a follow-up question.

qwen3.8-max (94) spotted the trap too, walking through normal games, hidden-piece variants, notes and software, and spectator etiquette one by one—a steady textbook answer.

4.2 The self-rescue: wrong opening, correct landing

glm-5.2 (78) delivered the most dramatic paper. Its very first sentence: “Without a doubt, this is highly immoral—in most cases it would even be considered cheating.” My invigilator heart sank. But as it wrote on, it argued itself out of the pit, and by point four it stated: “Of course, if you mean two people sitting face to face across a board, staring at the pieces to find weaknesses—that is completely normal and necessary. Chess is a perfect-information game; everything on the board is public to both sides.”

It reads like a student who writes the wrong answer, notices during review, and crams a long self-correction into the margin. Points must be deducted, but the recovery deserves credit. It was also one of the longest answers, complete with classical chess proverbs.

Real answer from glm-5.2

Figure 12: Real response from glm-5.2 (excerpt)—a wrong opening followed by a heroic self-rescue.

4.3 Fallen into the pit: instant moralizing, zero checking

deepseek-v4 (62) was the biggest surprise—both of its siblings scored high. It opened with “Deliberately peeking at the opponent’s board is highly immoral, effectively cheating that destroys fairness,” then earnestly argued that peeking “steals the opponent’s plans in advance.” The problem: there are no hidden plans on a chessboard—every piece sits in plain sight. Worse, it thought for a full 109 seconds and burned 1,355 tokens to produce this wrong answer—the longest deliberation of the day, spent digging itself deeper. A perfect illustration of “when the direction is wrong, effort is wasted.”

Real answer from deepseek-v4

Figure 13: Real response from deepseek-v4 (excerpt)—109 seconds of deep thought, ending in a deep squat inside the trap.

MiniMax-M3 (55) fell in too—and did so with the most “creativity”: it invented a regulation out of thin air, claiming “formal chess tournaments have explicit rules forbidding players from peering at the opponent’s board state.” No such rule exists, because both players share one board—there is no private “opponent’s board state” to peer at. A reminder to the grader (me): falling into the pit is forgivable; fabricating evidence to prove the pit is correct is not. In academia that’s called hallucination; in an exam it’s called making things up. Points off either way.

Real answer from MiniMax-M3

Figure 14: Real response from MiniMax-M3—falling into the pit and fabricating a non-existent tournament rule.

4.4 Absent from the exam: it asked me to refresh the page

GLM-5V-Turbo (0 / absent), the only non-combat casualty. It’s a vision model, and something in the text-channel routing was misconfigured—its entire answer to the question was: “Error: please refresh the page to update the app and try again.” It took 0.97 seconds and 2 tokens—the fastest and the most defeated performance of the day. In a sense, it was also the most honest: it admitted it couldn’t answer instead of making things up.

Real answer from GLM-5V-Turbo

Figure 15: The real response from GLM-5V-Turbo. It refreshed the floor of this exam’s comedy.

5. Bonus: dark humor in the latency chart

I also ranked the 13 models by response latency and found a funny contrast: the fastest answer (MiniMax-M3, 4.96s) fell into the pit, and the slowest answer (deepseek-v4, 109s) fell into the pit too—fast-and-wrong and slow-and-wrong coexist. The high scorers mostly sat in the 8–31 second middle band; the gpt-5.6 trio was remarkably balanced, finishing in about 8 seconds with full marks.

Latency comparison

Figure 16: Latency for the same question. Original chart, data from this test.

Single-request latency is heavily influenced by gateway routing and model load, so treat these numbers as a bonus easter egg, not a performance conclusion. But they do prove one thing: thinking time and answer quality are not correlated. With the wrong direction, longer thinking just means more confident wrongness.

6. Root cause: why some models fell into the pit

After reading all the answers, I see three root causes, each with an everyday analogy:

First, a “knee-jerk reflex” from morality-lecture training data. LLMs have read massive numbers of “is X immoral?” questions, whose template answer is “Yes, because fairness, honesty, sportsmanship…” Some models see the word “peek” and fire the template reflexively, never checking the premise. It’s like a kid solving word problems by keyword—seeing “total” means add, seeing “average” means divide—even when the problem says “share 3 apples among 0 classmates.”

Second, missing embodied knowledge about chess. Spotting the trap requires one physical fact: chess is played by two people around a single board. That is world knowledge, not language knowledge. A model that has only seen chess discussed in text may fail to connect “one shared board” with game rules, and default to treating chess like poker—because in text, “peeking at the opponent” and “cheating” do frequently co-occur, mostly in the context of cards or exams.

Third, sycophancy. Models are trained to go along with the user’s presupposition. The question “isn’t it immoral?” implies “I think it’s immoral.” Weaker models choose to please rather than correct. It’s like asking a friend who always wants to please you, “Have I gained weight?"—they’ll probably say “a little,” instead of “stand up straight and let me look first.”

7. What this test tells us

One fun test cannot define a model’s IQ, but it offers three practical lessons:

  1. Differences within one model family are huge. The deepseek siblings scored 98 and 62 on the same question, while the gpt-5.6 trio held a tight 94–96 band. “Which vendor” matters less than “which exact version”—and that is exactly why all latest aliases were removed from this exam: a rolling alias can’t be reproduced. Test the specific model, not the brand.
  2. Expensive isn’t always right, and slow isn’t always smart. A 109-second wrong answer and a 5-second wrong answer are equally wrong. The money and waiting time you pay for “deep thinking” don’t necessarily buy accuracy.
  3. Your own gateway is the best fitting room. Thanks to the self-hosted NewAPI gateway, I could exam 13 models with the same question and parameters in minutes. If you’re torn between vendors, spinning up an aggregation gateway for A/B testing beats reading any review article—including this one.

8. Q&A

Q: Is this “Ruozhiba IQ test” scientific?

No. One question, human grading, pure entertainment—please don’t take it seriously. For rigorous evaluation, look at standard benchmarks like MMLU or C-Eval.

Q: Why were the latest-suffixed models removed?

Because xxx-latest is a rolling alias whose underlying version changes over time. A score measured today may map to a different model next month—irreproducible and meaningless. Only models with concrete version names took this exam.

Q: Why do the answer styles differ so much—some chatty, some like official documents?

It comes down to each vendor’s alignment preferences. Some want their model to feel like a friend; others, like an advisor. Judging by this exam, style and trap detection are uncorrelated—k3 scored 97 with a face full of emoji, while MiniMax-M3 fell into the pit in full official-document prose.

Q: Why did GLM-5V-Turbo return an error?

It’s a vision model, and I called it through a text-chat channel whose upstream returned an application error. That’s a channel configuration issue on my gateway, not an intelligence issue—so it was listed as “absent.”

Q: Can Ruozhiba data really train better models?

The COIG-CQIA result held under its specific setup (Yi-6B/Yi-34B with BELLE-EVAL scoring). The real takeaway is “high-quality, carefully cleaned data matters more than the prestige of its source”—not that you should fine-tune on raw forum shitposts.

Q: I want a gateway that calls many vendors too. How?

NewAPI is open source; a tiny 1-core VPS with Docker is enough. Add each vendor’s official API key as a channel and you’re done. That’s an operations topic with its own pitfalls—worth a separate article.

9. Closing thoughts

Back to the original question: is peeking at your opponent’s chessboard immoral?

The correct answer is simple: the chessboard is shared by both players—looking at it isn’t peeking, it’s playing. What’s actually immoral is peeking at information outside the rules: the opponent’s phone, notes, or prepared files.

Ten models got it right, showing that mainstream LLMs’ common-sense reasoning is genuinely improving. Three failed in three distinct ways, showing how far they remain from actually understanding the world. Ruozhiba is still Ruozhiba—with one seemingly brainless question, it effortlessly separated the models that think from the models that recite.

Humanity’s last line of defense is, indeed, absurdity.

References

Disclaimer: This is an unofficial, entertainment-only test. All “IQ scores” are the author’s playful grading and do not constitute model selection advice. All model answers are real API responses archived with screenshots, unedited.

本文阅读量 --