<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0"
     xmlns:atom="http://www.w3.org/2005/Atom"
     xmlns:content="http://purl.org/rss/1.0/modules/content/"
     xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>Pavlo Kostenko</title>
    <link>https://pavlokostenko.com/</link>
    <description>Notes on marketing attribution, martech engineering, and AI in analytics.</description>
    <language>en</language>
    <lastBuildDate>Sun, 26 Jul 2026 00:00:00 GMT</lastBuildDate>
    <atom:link href="https://pavlokostenko.com/rss.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The Crowbar Problem: One Riddle, 600 AI Answers, and the Paragraph That Flipped Every Conclusion</title>
      <link>https://pavlokostenko.com/articles/crowbar-problem/</link>
      <guid isPermaLink="true">https://pavlokostenko.com/articles/crowbar-problem/</guid>
      <pubDate>Sun, 26 Jul 2026 00:00:00 GMT</pubDate>
      <dc:creator>Pavlo Kostenko</dc:creator>
      <category>AI in Analytics</category>
      <description>Five Claude models answered one lateral thinking riddle 600 times, across eight languages and five reasoning effort levels. Adding a single paragraph to the judges&apos; instructions moved the full solution rate from 83% to 24% without changing a byte of the data.</description>
      <content:encoded><![CDATA[<p>Five Claude models answered one lateral thinking riddle 600 times, across eight languages and five reasoning effort levels. Adding a single paragraph to the judges' instructions moved the full solution rate from 83% to 24% without changing a byte of the data.</p>
<ol><li>In a controlled test of 600 answers to one lateral thinking riddle, Claude Haiku 4.5 refused to answer 87% of the time in Ukrainian and 0% of the time in English.</li><li>Across the same 600 answers, Fable 5 found the full solution 83 times out of 120, Opus 5 found it 37 times, Opus 4.8 found it 25 times, and Sonnet 5 and Haiku 4.5 found it zero times.</li><li>Sonnet 5 produced a physically correct partial answer in all 120 attempts and never once used the tool the riddle handed it, which is the difference between a right answer and the right answer.</li><li>Raising reasoning effort from low to high took full solutions from 14 to 35 out of 120 per level, and the two effort levels above high added three more between them.</li><li>Adding one paragraph defining the correct answer to the judges' instructions moved the overall full solution rate from 83% to 24%, with no change to the underlying answers.</li><li>That same paragraph moved Sonnet 5 from an apparent 1.98 out of 2 to a flat 1.00, and turned reasoning effort from a waste of budget into a lever that nearly tripled results.</li><li>An LLM judge given no reference answer grades confidence and fluency rather than correctness, which is why reference free evaluation runs systematically generous.</li><li>The entire experiment cost $31.32 in compute and roughly two hours of work, which is less than most teams spend deciding which public leaderboard to believe.</li><li>Opus 5 at maximum reasoning effort took a median of 95 seconds per answer and cost about $0.14, roughly 27 times the price of a Haiku 4.5 answer.</li><li>One riddle is a probe rather than a benchmark, so the conclusions about evaluation method travel much further than the conclusions about any individual model.</li></ol>
<p><a href="https://pavlokostenko.com/articles/crowbar-problem/">Read the full piece</a></p>]]></content:encoded>
    </item>
  </channel>
</rss>
