Introduction
In my previous deep dive on MathCanvas, my conclusion was: Large models are unstable in geometry not because they cannot read diagrams, but because they lack a stable intermediate structure to operate on.
Some research has started to let models draw first, then reason.
For example, MathCanvas has the model generate internal sketches and then reason over those sketches.
After that article, a reader asked:
If visual intermediate steps are so important, why not let the model actually draw the figure?
I looked for related work and found exactly that in HKU's CodePlot-CoT.

The model no longer "imagines" auxiliary lines. It writes Python matplotlib code, draws the lines explicitly, and continues solving.
This sounds very reasonable: if visual reasoning is unstable, provide an executable visual world.
But a new question appears:
When a model starts writing plotting code, is it doing geometric reasoning, or only numerical validation on one concrete coordinate instance?
To answer that, we need to clarify what problem this paper actually solves.
What Problem the Paper Actually Solves
CodePlot-CoT focuses on a lower-level phenomenon: unstable spatial working memory in multimodal models for math problems.
Concretely, the model can understand the prompt and produce a standard chain of thought, but once intermediate geometric states are required, reasoning drifts.
This shows up as:
- inconsistent auxiliary lines across steps,
- forgotten spatial relations,
- later steps relying on structures that do not actually exist.
MathCanvas addressed this by generating internal sketches as visual CoT (Visual Chain-of-Thought).
CodePlot-CoT takes another route: instead of asking the model to imagine a figure, let it operate in a real executable plotting environment.
In other words, the "figure in thought" is outsourced to Python.
Technical Core: Let the Model Write Matplotlib

In one paper example, the model produces a step:
"Connect C and D first"
Instead of continuing in text, it outputs Python code directly:
ax.plot([C[0], D[0]], [C[1], D[1]])
The full loop is:
text reasoning -> plotting code generation -> image rendering -> model re-input -> continued reasoning
So the model's intermediate state no longer lives only in tokens or latent space; it lives in an executable external environment.
This brings immediate benefits:
1. Stable spatial state
The model depends on the environment instead of memory (similar to tool-using agents).
2. Better visual consistency
Common multimodal "diagram drift" is reduced.
3. Scalable data generation
The paper builds Math-VR (178k problems), turning "figure -> code -> reasoning" into supervised signal.
This is a classic computer-vision style move: do not force the model to imagine the world; train it with an operational world.
Up to this point, the paper is elegant. But there is a direct limitation.
Key Question: Does It Really Understand Geometry?
Look again at this line:
ax.plot([C[0], D[0]], [C[1], D[1]])
It means: draw a segment in a coordinate plane.
But geometry reasoning needs more than that.
In geometry, "connect C and D" is not just drawing. It is a construction operation:
- Is CD a chord?
- Is CD an angle bisector?
- Is CD perpendicular to a given line?
- Is CD the line through intersections of two circles?
These are where reasoning power comes from.
Matplotlib expresses visual appearance. Geometry reasoning needs relational constraints.
So the intermediate structure is still "looks right," not "must be true."
Loss of Construction Causality
Geometric proof depends on "why true," not merely "looks true."
For example, construction: angle bisector -> equal angles is a logical implication.
In rendered-image reasoning, this becomes measured angles are approximately equal.
These are fundamentally different:
| Type | Nature |
|---|---|
| Geometric construction | Necessary |
| Numeric instance | Contingent |
CodePlot-CoT reasoning is:
Generate one coordinate instance -> inspect image -> conclude
Mathematically, this is single-instance validation. But geometric statements require validity across all configurations, while the model sees only one sampled world.
Visual Validation vs Mathematical Proof
At this point we can abstract two paradigms:
| CodePlot-CoT | Geometric proof | |
|---|---|---|
| Basis of judgment | Looks valid | Necessarily valid |
| Method | Experiment | Derivation |
| Nature | Empirical reasoning | Deductive reasoning |
CodePlot-CoT is effectively a "geometry experiment AI," not a "geometry proof AI."
It answers: "Does this image support the claim?" rather than "Is this claim valid in an axiomatic system?"
Why It Does Not Reach AlphaGeometry
HKU CodePlot-CoT provides executable graphics; Google AlphaGeometry provides derivable proofs.
What is missing in between is the geometric object layer itself.
Concretely:
- Is a point an intersection point or a free point?
- Is a line an angle bisector or just a connector?
- Is a circle uniquely defined by three points, or arbitrarily drawn?
This is not a visual problem. It is a mathematical structure-modeling problem.
If we place the routes on one axis:
image -> rendering code -> geometric objects -> logical proof
CodePlot-CoT stops at layer 2, AlphaGeometry is at layer 4, and layer 3 is where humans spend most of their solving effort: geometric construction.
When humans solve geometry, we rarely start with formal proofs, and we do not rely only on images. We operate on objects: draw perpendiculars, take intersections, construct circles through points, construct angle bisectors.
This step is neither pure vision nor final proof, but it determines whether all subsequent reasoning can hold.

Final Note
In my own project, Dino-GSP (大角几何), I am trying to isolate exactly this missing layer.
The goal is neither "let the model draw" nor "force direct formal proof," but let the model operate on geometric objects themselves.
The model output is not ax.plot(...), and not Therefore AB ⟂ CD, but operations such as:
PerpLine(<Line>, <Point>) # Construct a perpendicular through an external point
Intersection(Circle(O, 2), Line(2, 3)) # Get all intersections of a circle and a line
Once the intermediate representation becomes object constraints, many things change:
- figures can be generated stably instead of relying on one coordinate instance,
- relations can be directly verified instead of visually measured,
- reasoning can form a formal chain instead of an empirical guess.
From this perspective, CodePlot-CoT matters not because it is simply "stronger at solving math," but because it proves that visual reasoning needs an external intermediate workspace.
That is, the LLM bottleneck is not "it cannot reason mathematically," but "it lacks a workspace for constructing mathematical models."
CodePlot-CoT provides one kind of workspace, but it is not yet a mathematical workspace.
What geometry may truly need is an operable geometric language.