HKU CodePlot-CoT Deep Dive: Visual Reasoning or Geometric Reasoning?

Theme: AI for Math

Published:

Last Updated:

Introduction

In my previous deep dive on MathCanvas, my conclusion was: Large models are unstable in geometry not because they cannot read diagrams, but because they lack a stable intermediate structure to operate on.

Some research has started to let models draw first, then reason.

For example, MathCanvas has the model generate internal sketches and then reason over those sketches.

After that article, a reader asked:

If visual intermediate steps are so important, why not let the model actually draw the figure?

I looked for related work and found exactly that in HKU's CodePlot-CoT.

CodePlot-CoT example

The model no longer "imagines" auxiliary lines. It writes Python matplotlib code, draws the lines explicitly, and continues solving.

This sounds very reasonable: if visual reasoning is unstable, provide an executable visual world.

But a new question appears:

When a model starts writing plotting code, is it doing geometric reasoning, or only numerical validation on one concrete coordinate instance?

To answer that, we need to clarify what problem this paper actually solves.

What Problem the Paper Actually Solves

CodePlot-CoT focuses on a lower-level phenomenon: unstable spatial working memory in multimodal models for math problems.

Concretely, the model can understand the prompt and produce a standard chain of thought, but once intermediate geometric states are required, reasoning drifts.

This shows up as:

  • inconsistent auxiliary lines across steps,
  • forgotten spatial relations,
  • later steps relying on structures that do not actually exist.

MathCanvas addressed this by generating internal sketches as visual CoT (Visual Chain-of-Thought).

CodePlot-CoT takes another route: instead of asking the model to imagine a figure, let it operate in a real executable plotting environment.

In other words, the "figure in thought" is outsourced to Python.

Technical Core: Let the Model Write Matplotlib

CodePlot-CoT pipeline

In one paper example, the model produces a step:

"Connect C and D first"

Instead of continuing in text, it outputs Python code directly:

ax.plot([C[0], D[0]], [C[1], D[1]])

The full loop is:

text reasoning -> plotting code generation -> image rendering -> model re-input -> continued reasoning

So the model's intermediate state no longer lives only in tokens or latent space; it lives in an executable external environment.

This brings immediate benefits:

1. Stable spatial state

The model depends on the environment instead of memory (similar to tool-using agents).

2. Better visual consistency

Common multimodal "diagram drift" is reduced.

3. Scalable data generation

The paper builds Math-VR (178k problems), turning "figure -> code -> reasoning" into supervised signal.

This is a classic computer-vision style move: do not force the model to imagine the world; train it with an operational world.

Up to this point, the paper is elegant. But there is a direct limitation.

Key Question: Does It Really Understand Geometry?

Look again at this line:

ax.plot([C[0], D[0]], [C[1], D[1]])

It means: draw a segment in a coordinate plane.

But geometry reasoning needs more than that.

In geometry, "connect C and D" is not just drawing. It is a construction operation:

  • Is CD a chord?
  • Is CD an angle bisector?
  • Is CD perpendicular to a given line?
  • Is CD the line through intersections of two circles?

These are where reasoning power comes from.

Matplotlib expresses visual appearance. Geometry reasoning needs relational constraints.

So the intermediate structure is still "looks right," not "must be true."

Loss of Construction Causality

Geometric proof depends on "why true," not merely "looks true."

For example, construction: angle bisector -> equal angles is a logical implication.

In rendered-image reasoning, this becomes measured angles are approximately equal.

These are fundamentally different:

TypeNature
Geometric constructionNecessary
Numeric instanceContingent

CodePlot-CoT reasoning is:

Generate one coordinate instance -> inspect image -> conclude

Mathematically, this is single-instance validation. But geometric statements require validity across all configurations, while the model sees only one sampled world.

Visual Validation vs Mathematical Proof

At this point we can abstract two paradigms:

CodePlot-CoTGeometric proof
Basis of judgmentLooks validNecessarily valid
MethodExperimentDerivation
NatureEmpirical reasoningDeductive reasoning

CodePlot-CoT is effectively a "geometry experiment AI," not a "geometry proof AI."

It answers: "Does this image support the claim?" rather than "Is this claim valid in an axiomatic system?"

Why It Does Not Reach AlphaGeometry

HKU CodePlot-CoT provides executable graphics; Google AlphaGeometry provides derivable proofs.

What is missing in between is the geometric object layer itself.

Concretely:

  • Is a point an intersection point or a free point?
  • Is a line an angle bisector or just a connector?
  • Is a circle uniquely defined by three points, or arbitrarily drawn?

This is not a visual problem. It is a mathematical structure-modeling problem.

If we place the routes on one axis:

image -> rendering code -> geometric objects -> logical proof

CodePlot-CoT stops at layer 2, AlphaGeometry is at layer 4, and layer 3 is where humans spend most of their solving effort: geometric construction.

When humans solve geometry, we rarely start with formal proofs, and we do not rely only on images. We operate on objects: draw perpendiculars, take intersections, construct circles through points, construct angle bisectors.

This step is neither pure vision nor final proof, but it determines whether all subsequent reasoning can hold.

Reasoning layer diagram

Final Note

In my own project, Dino-GSP (大角几何), I am trying to isolate exactly this missing layer.

The goal is neither "let the model draw" nor "force direct formal proof," but let the model operate on geometric objects themselves.

The model output is not ax.plot(...), and not Therefore AB ⟂ CD, but operations such as:

PerpLine(<Line>, <Point>)  # Construct a perpendicular through an external point
Intersection(Circle(O, 2), Line(2, 3))  # Get all intersections of a circle and a line

Once the intermediate representation becomes object constraints, many things change:

  • figures can be generated stably instead of relying on one coordinate instance,
  • relations can be directly verified instead of visually measured,
  • reasoning can form a formal chain instead of an empirical guess.

From this perspective, CodePlot-CoT matters not because it is simply "stronger at solving math," but because it proves that visual reasoning needs an external intermediate workspace.

That is, the LLM bottleneck is not "it cannot reason mathematically," but "it lacks a workspace for constructing mathematical models."

CodePlot-CoT provides one kind of workspace, but it is not yet a mathematical workspace.

What geometry may truly need is an operable geometric language.

References