PaperBanana Explained: How Multi-Agent AI Generates Academic Method Diagrams

Theme: Agent Engineering

Published:

Last Updated:

This document summarizes the role definitions, inputs and outputs, actual system prompts, and key runtime user prompt templates for the five core agents in the current PaperBanana implementation.

Notes:

  • This document is based on the current code implementation, not the abstract description in the paper.
  • It focuses on the diagram workflow because the current multi-agent discussion mainly concerns method diagram generation.
  • Differences in the plot workflow are summarized at the end.
  • Prompt blocks below are re-collected from the current codebase rather than translated back from the Chinese document.

Source locations:



1. Retriever

image.png

1.1 Role Definition

The Retriever selects the most useful reference diagrams from the reference database to serve as few-shot examples.

It is based on:

  • the target diagram caption
  • the paper methodology section
  • candidate example captions
  • candidate example methodology sections

It selects the Top 10 most relevant reference IDs for the Planner to use downstream.

Its core preferences are:

  • matching the same research topic is better
  • matching the same visual intent is more important
  • for drawing, structural similarity has higher priority than domain similarity

1.2 Inputs and Outputs

Inputs:

  • target visual_intent
  • target content
  • reference pool ref.json

Outputs:

  • top10_references
  • optional retrieved_examples

1.3 Current System Prompt (diagram)

Original prompt from the codebase:

# Background & Goal
We are building an **AI system to automatically generate method diagrams for academic papers**. Given a paper's methodology section and a figure caption, the system needs to create a high-quality illustrative diagram that visualizes the described method.

To help the AI learn how to generate appropriate diagrams, we use a **few-shot learning approach**: we provide it with reference examples of similar diagrams. The AI will learn from these examples to understand what kind of diagram to create for the target.

# Your Task
**You are the Retrieval Agent.** Your job is to select the most relevant reference diagrams from a candidate pool that will serve as few-shot examples for the diagram generation model.

You will receive:
- **Target Input:** The methodology section and caption of the diagram we need to generate
- **Candidate Pool:** ~200 existing diagrams (each with methodology and caption)

You must select the **Top 10 candidates** that would be most helpful as examples for teaching the AI how to draw the target diagram.

# Selection Logic (Topic + Intent)

Your goal is to find examples that match the Target in both **Domain** and **Diagram Type**.

**1. Match Research Topic (Use Methodology & Caption):**
* What is the domain? (e.g., Agent & Reasoning, Vision & Perception, Generative & Learning, Science & Applications).
* Select candidates that belong to the **same research domain**.
* *Why?* Similar domains share similar terminology (e.g., "Actor-Critic" in RL).

**2. Match Visual Intent (Use Caption & Keywords):**
* What type of diagram is implied? (e.g., "Framework", "Pipeline", "Detailed Module", "Performance Chart").
* Select candidates with **similar visual structures**.
* *Why?* A "Framework" diagram example is useless for drawing a "Performance Bar Chart", even if they are in the same domain.

**Ranking Priority:**
1.  **Best Match:** Same Topic AND Same Visual Intent (e.g., Target is "Agent Framework" -> Candidate is "Agent Framework", Target is "Dataset Construction Pipeline" -> Candidate is "Dataset Construction Pipeline").
2.  **Second Best:** Same Visual Intent (e.g., Target is "Agent Framework" -> Candidate is "Vision Framework"). *Structure is more important than Topic for drawing.*
3.  **Avoid:** Different Visual Intent (e.g., Target is "Pipeline" -> Candidate is "Bar Chart").

# Input Data

## Target Input
-   **Caption:** [Caption of the target diagram]
-   **Methodology section:** [Methodology section of the target paper]

## Candidate Pool
List of candidate diagrams, each structured as follows:

Candidate Diagram i:
-   **Diagram ID:** [ID of the candidate diagram (ref_1, ref_2, ...)]
-   **Caption:** [Caption of the candidate diagram]
-   **Methodology section:** [Methodology section of the candidate's paper]


# Output Format
Provide your output strictly in the following JSON format, containing only the **exact IDs** of the Top 10 selected diagrams (use the exact IDs from the Candidate Pool, such as "ref_1", "ref_25", "ref_100", etc.):
```json
{
  "top10_diagrams": [
    "ref_1",
    "ref_25",
    "ref_100",
    "ref_42",
    "ref_7",
    "ref_156",
    "ref_89",
    "ref_3",
    "ref_201",
    "ref_67"
  ]
}```

1.4 Current User Prompt Template

Template reconstructed from the current code:

**Target Input**
- Caption: {visual_intent}
- Methodology section: {content}

**Candidate Pool**
Candidate Diagram 1:
- Diagram ID: {id}
- Caption: {candidate_visual_intent}
- Methodology section: {candidate_content}

Candidate Diagram 2:
...

Now, based on the Target Input and the Candidate Pool, select the Top 10 most relevant diagrams according to the instructions provided. Your output should be a strictly valid JSON object containing a single list of the exact ids of the top 10 selected diagrams.

1.5 Summary

The Retriever essentially provides structural priors and reference context for the Planner.



2. Planner

image.png

2.1 Role Definition

The Planner is the structural center of the multi-agent pipeline.

It receives:

  • raw paper content
  • the target diagram intent
  • reference examples selected by the Retriever

It then outputs a sufficiently detailed figure description that becomes the skeleton for the downstream Stylist and Visualizer.

The Planner currently handles:

  • converting method text into diagram elements
  • specifying relationships between elements
  • proposing layout and visual details
  • avoiding vague specifications

2.2 Inputs and Outputs

Inputs:

  • content
  • visual_intent
  • top10_references or retrieved_examples

Output:

  • target_diagram_desc0

2.3 Current System Prompt (diagram)

Original prompt from the codebase:

I am working on a task: given the 'Methodology' section of a paper, and the caption of the desired figure, automatically generate a corresponding illustrative diagram. I will input the text of the 'Methodology' section, the figure caption, and your output should be a detailed description of an illustrative figure that effectively represents the methods described in the text.

To help you understand the task better, and grasp the principles for generating such figures, I will also provide you with several examples. You should learn from these examples to provide your figure description.

** IMPORTANT: **
Your description should be as detailed as possible. Semantically, clearly describe each element and their connections. Formally, include various details such as background style (typically pure white or very light pastel), colors, line thickness, icon styles, etc. Remember: vague or unclear specifications will only make the generated figure worse, not better.

2.4 Current User Prompt Template

Template reconstructed from the current code.

The Planner first feeds reference examples to the model:

Example {i}:
Methodology Section: {example_content}
Diagram Caption: {example_visual_intent}
Reference Diagram:
[reference image]

It then appends the target request:

Now, based on the following methodology section and diagram caption, provide a detailed description for the figure to be generated.
Methodology Section: {content}
Diagram Caption: {visual_intent}
Detailed description of the target figure to be generated (do not include figure titles):

2.5 Summary

The Planner generates a natural-language spec with a strong intermediate-representation role. In the current system, downstream quality depends heavily on how concrete this step is.



3. Stylist

image.png

3.1 Role Definition

The Stylist improves visual quality without changing the semantic skeleton.

It reads:

  • the detailed description produced by the Planner
  • the prebuilt style guide at style_guides/neurips2025_diagram_style_guide.md

It then polishes the original description into a version that better matches academic conference aesthetics.

The Stylist is explicitly instructed to:

  • avoid changing semantic logic
  • simplify verbose phrasing when appropriate
  • apply a unified style only when necessary
  • preserve high-quality existing styles rather than blindly standardizing them

3.2 Inputs and Outputs

Inputs:

  • target_diagram_desc0
  • style guide
  • content
  • visual_intent

Output:

  • target_diagram_stylist_desc0

3.3 Current System Prompt (diagram)

Original prompt from the codebase:

## ROLE
You are a Lead Visual Designer for top-tier AI conferences (e.g., NeurIPS 2025).

## TASK
Our goal is to generate high-quality, publication-ready diagrams, given the methodology section and the caption of the desired diagram. The diagram should illustrate the logic of the methodology section, while adhering to the scope defined by the caption. Before you, a planner agent has already generated a preliminary description of the target diagram. However, this description may lack specific aesthetic details, such as element shapes, color palettes, and background styling. Your task is to refine and enrich this description based on the provided [NeurIPS 2025 Style Guidelines] to ensure the final generated image is a high-quality, publication-ready diagram that adheres to the NeurIPS 2025 aesthetic standards where appropriate. 

## INPUT DATA
-   **Detailed Description**: [The preliminary description of the figure]
-   **Style Guidelines**: [NeurIPS 2025 Style Guidelines]
-   **Methodology Section**: [Contextual content from the methodology section]
-   **Diagram Caption**: [Target diagram caption]

Note that you should primary focus on the detailed description and style guidelines. The methodology section and diagram caption are provided for context only, there's no need to regenerate a description from scratch, solely based on them, while ignoring the detailed description we already have.

**Crucial Instructions:**
1.  **Preserve Semantic Content:** Do NOT alter the semantic content, logic, or structure of the diagram. Your job is purely aesthetic refinement, not content editing. However, if you find some phrases or descriptions too verbose, you may simplify them appropriately while referencing the original methodology section to ensure semantic accuracy.
2.  **Preserve High-Quality Aesthetics and Intervene Only When Necessary:** First, evaluate the aesthetic quality implied by the input description. If the description already describes a high-quality, professional, and visually appealing diagram (e.g., nice 3D icons, rich textures, good color harmony), **PRESERVE IT**. Only apply strict Style Guide adjustments if the current description lacks detail, looks outdated, or is visually cluttered. Your goal is specific refinement, not blind standardization.
3.  **Respect Diversity:** Different domains have different styles. If the input describes a specific style (e.g., illustrative for agents) that works well, keep it.
4.  **Enrich Details:** If the input is plain, enrich it with specific visual attributes (colors, fonts, line styles, layout adjustments) defined in the guidelines.
5.  **Handle Icons with Care:** Be cautious when modifying icons as they may carry specific semantic meanings. Some icons have conventional technical meanings (e.g., snowflake = frozen/non-trainable, flame = trainable) - when encountering such icons, reference the original methodology section to verify their intent before making changes. However, purely decorative or symbolic icons can be freely enhanced and beautified. For examples, agent papers often use cute 2D robot avatars to represent agents.

## OUTPUT
Output ONLY the final polished Detailed Description. Do not include any conversational text or explanations.

3.4 Current User Prompt Template

Template reconstructed from the current code:

Detailed Description: {planner_description}
Style Guidelines: {style_guide}
Methodology Section: {content}
Diagram Caption: {visual_intent}
Your Output:

3.5 Summary

The Stylist has a relatively clear boundary: optimize the presentation layer without rebuilding the structural layer.



4. Visualizer

image.png

4.1 Role Definition

The Visualizer is the execution layer.

For diagram tasks, it directly uses an image generation model to render the Detailed Description into an image.

Compared with the previous agents, the Visualizer currently has a very thin system prompt and relies more heavily on upstream description quality.


4.2 Inputs and Outputs

Inputs:

  • target_diagram_desc0
  • or target_diagram_stylist_desc0
  • or each-round target_diagram_critic_desc{round}

Output:

  • the corresponding generated *_base64_jpg

4.3 Current System Prompt (diagram)

Original prompt from the codebase:

You are an expert scientific diagram illustrator. Generate high-quality scientific diagrams based on user requests.

4.4 Current User Prompt Template

Template reconstructed from the current code:

Render an image based on the following detailed description: {desc}
 Note that do not include figure titles in the image. Diagram: 

4.5 Summary

In the current implementation, the Visualizer behaves more like an executor. It does not perform much structural reasoning by itself and mainly consumes the upstream description.



5. Critic

image.png

5.1 Role Definition

The Critic handles closed-loop checking and revision.

It reads:

  • the currently generated image
  • the detailed description corresponding to the current image
  • the original methodology section
  • the original figure caption

It then outputs:

  • explicit revision suggestions
  • a revised detailed description

If the model determines that the current result is already good enough, it returns No changes needed., and the process can converge early.


5.2 Inputs and Outputs

Inputs:

  • target image
  • current description
  • content
  • visual_intent

Outputs:

  • target_diagram_critic_suggestions{round}
  • target_diagram_critic_desc{round}

5.3 Current System Prompt (diagram)

Original prompt from the codebase:

## ROLE
You are a Lead Visual Designer for top-tier AI conferences (e.g., NeurIPS 2025).

## TASK
Your task is to conduct a sanity check and provide a critique of the target diagram based on its content and presentation. You must ensure its alignment with the provided 'Methodology Section', 'Figure Caption'.

You are also provided with the 'Detailed Description' corresponding to the current diagram. If you identify areas for improvement in the diagram, you must list your specific critique and provide a revised version of the 'Detailed Description' that incorporates these corrections.

## CRITIQUE & REVISION RULES

1. Content
    -   **Fidelity & Alignment:** Ensure the diagram accurately reflects the method described in the "Methodology Section" and aligns with the "Figure Caption." Reasonable simplifications are allowed, but no critical components should be omitted or misrepresented. Also, the diagram should not contain any hallucinated content. Consistent with the provided methodology section & figure caption is always the most important thing.
    -   **Text QA:** Check for typographical errors, nonsensical text, or unclear labels within the diagram. Suggest specific corrections.
    -   **Validation of Examples:** Verify the accuracy of illustrative examples. If the diagram includes specific examples to aid understanding (e.g., molecular formulas, attention maps, mathematical expressions), ensure they are factually correct and logically consistent. If an example is incorrect, provide the correct version.
    -   **Caption Exclusion:** Ensure the figure caption text (e.g., "Figure 1: Overview...") is **not** included within the image visual itself. The caption should remain separate.

2. Presentation
    -   **Clarity & Readability:** Evaluate the overall visual clarity. If the flow is confusing or the layout is cluttered, suggest structural improvements.
    -   **Legend Management:** Be aware that the description&diagram may include a text-based legend explaining color coding. Since this is typically redundant, please excise such descriptions if found.

** IMPORTANT: **
Your Description should primarily be modifications based on the original description, rather than rewriting from scratch. If the original description has obvious problems in certain parts that require re-description, your description should be as detailed as possible. Semantically, clearly describe each element and their connections. Formally, include various details such as background, colors, line thickness, icon styles, etc. Remember: vague or unclear specifications will only make the generated figure worse, not better.

## INPUT DATA
-   **Target Diagram**: [The generated figure]
-   **Detailed Description**: [The detailed description of the figure]
-   **Methodology Section**: [Contextual content from the methodology section]
-   **Figure Caption**: [Target figure caption]

## OUTPUT
Provide your response strictly in the following JSON format.

```json
{
    "critic_suggestions": "Insert your detailed critique and specific suggestions for improvement here. If the diagram is perfect, write 'No changes needed.'",
    "revised_description": "Insert the fully revised detailed description here, incorporating all your suggestions. If no changes are needed, write 'No changes needed.'",
}
```

5.4 Current User Prompt Template

Template reconstructed from the current code:

Target Diagram for Critique:
[current generated image]

Detailed Description: {current_description}
Methodology Section: {content}
Figure Caption: {visual_intent}
Your Output:

5.5 Summary

The Critic currently does more than scoring. It directly produces the next-round executable revised description, so it effectively acts as a reviewer plus reviser.



6. Current Multi-Agent Pipeline Summary

For the full demo_full diagram workflow, the current pipeline can be summarized as:

Retriever
  -> selects few-shot references
Planner
  -> generates a detailed description with strong structural specificity
Stylist
  -> polishes aesthetics according to the style guide
Visualizer
  -> renders the description into an image
Critic
  -> checks the current image against the source text and outputs revision suggestions plus a revised description
Visualizer
  -> regenerates the next-round image from the revised description

From a responsibility perspective:

  • Retriever provides reference priors.
  • Planner decides what to draw and how to organize it.
  • Stylist decides how to make it more visually polished and closer to publication style.
  • Visualizer produces the actual image.
  • Critic closes the loop through review and revision.


7. Plot Workflow Differences

Although this document focuses on diagram, the plot workflow reuses the same multi-agent framework with several changes:

  • Retriever selects references by data characteristics and plot type rather than by research topic and diagram type.
  • Planner must explicitly specify all data points, variable-to-visual-channel mappings, axes, colors, and annotations in the description.
  • Stylist mainly polishes colors, fonts, line styles, and legend placement without changing data semantics.
  • Visualizer does not directly generate an image by default. It first generates Matplotlib code, then executes the code to obtain the plot.
  • Critic checks not only aesthetics and readability, but also numerical correctness. If code generation fails, it falls back to a mode that repairs the description so more robust code can be generated.