No description
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
2026-05-08 12:55:47 +02:00
framework fix: continue_from_checkpoint generation numbering bugs 2026-05-08 12:01:07 +02:00
problems Skip gen_0000 checkpoint on broken initial solution 2026-05-08 10:50:23 +02:00
templates Simplify to solution.py — always the current best 2026-05-08 10:44:49 +02:00
.gitignore Initial commit: Evolvium — AI-driven code evolution framework 2026-05-08 10:10:03 +02:00
evolve.py Simplify to solution.py — always the current best 2026-05-08 10:44:49 +02:00
README.md Improve flowchart: remove arrow symbols from labels, wrap SAVE/DISCARD in subgraph, add descriptive edge labels 2026-05-08 12:55:47 +02:00

Evolvium

AI-driven code evolution — the AI iteratively mutates a Python function to maximize a fitness score, keeping improvements and discarding regressions.

How it works

flowchart LR
    %% ── input files ──
    PM["📄 problem.md<br/>task description"]
    SOL["📄 solution.py<br/>current code"]
    EPY["📄 evaluation.py<br/>scoring logic"]

    %% ── core loop ──
    EVAL["🔧 Evaluate<br/>run code, compute score"]
    CMP{"Score improved?"}
    LLM[("🧠 LLM<br/>mutate code")]
    CKPT["📁 checkpoints/<br/>historical snapshots"]

    %% ── decision block: SAVE wraps DISCARD ──
    subgraph outcome["After comparison"]
        direction TB
        SAVE["💾 Save checkpoint<br/>Update solution.py"]
        DISCARD["🗑️ Discard<br/>(revert to previous)"]
    end

    PM -->|"reads problem description"| LLM
    SOL -->|"current code"| EVAL
    EPY -->|"scoring logic"| EVAL

    EVAL -->|"fitness score"| CMP
    CMP -->|"yes – keep mutation"| SAVE
    CMP -->|"no – revert"| DISCARD

    SAVE -->|"best code so far"| LLM
    DISCARD -->|"previous code\n(unchanged)"| LLM
    LLM -->|"mutated code"| EVAL

    SAVE -.->|"checkpoint snapshot"| CKPT

    %% ── styles ──
    classDef input fill:#dbeafe,stroke:#3b82f6,color:#1e3a5f
    classDef compute fill:#fef3c7,stroke:#f59e0b,color:#78350f
    classDef llm fill:#ede9fe,stroke:#8b5cf6,color:#4c1d95
    classDef success fill:#d1fae5,stroke:#10b981,color:#064e3b
    classDef discard fill:#fee2e2,stroke:#ef4444,color:#7f1d1d
    classDef storage fill:#f3f4f6,stroke:#6b7280,color:#1f2937

    class PM,SOL,EPY input
    class EVAL,CMP compute
    class LLM llm
    class SAVE success
    class DISCARD discard
    class CKPT storage
  1. Read solution.py — the current best code (inside # EVOLVE-BLOCK-START / END markers)
  2. Evaluate — run evaluation.py to compute a fitness score
  3. Mutate — send the problem description + current code + score to an LLM, asking it to improve the code
  4. Compare — if the mutated code scores higher, it becomes the new best; checkpoint saved, solution.py updated
  5. Repeat for N generations

If the initial solution.py is broken or missing, the LLM gets it anyway and can produce a valid first version from scratch.

Quick start

# Test the loop without an LLM (mock returns code unchanged)
python evolve.py --problem-dir sum_target --llm mock

# Run with a real LLM (requires pi)
python evolve.py --problem-dir sum_target --llm pi --generations 5

# Use a different model
python evolve.py --problem-dir tammes --llm pi --model deepseek/deepseek-v4-pro

Creating a new problem

Each problem needs exactly 3 files:

File Purpose
problem.md Plain-English description of the optimization task. The LLM reads this.
solution.py Python code with # EVOLVE-BLOCK-START / # EVOLVE-BLOCK-END markers. The LLM only mutates code between these markers.
evaluation.py Two functions: get_score(result) -> float (higher = better) and evaluate(data) -> (metrics, outputs). The harness calls get_positions(n, starting_positions) from solution.py.

Copy templates/PROBLEM_TEMPLATE.py as a starting point, or copy an existing problem directory:

cp -r problems/sum_target problems/my_problem
# Edit problem.md, evaluation.py, solution.py
python evolve.py --problem-dir my_problem --llm mock   # test
python evolve.py --problem-dir my_problem --llm pi     # run

LLM backends

Backend Flag Notes
pi --llm pi Calls pi -p --no-session --no-tools as a subprocess. Default model: deepseek/deepseek-v4-flash. Set EVOLVIUM_MODEL env var or use --model.
mock --llm mock Returns the input program unchanged. For testing the evolution loop.
openai --llm openai Uses the OpenAI API. Requires OPENAI_API_KEY. Set OPENAI_MODEL env var or use --model.

CLI reference

python evolve.py --problem-dir <name> [options]

Options:
  --problem-dir, -d   Problem directory name or path (required)
  --generations, -g   Number of generations (default: 10)
  --llm, -l           LLM backend: pi, mock, openai, auto (default: auto)
  --model             Override model (e.g. deepseek/deepseek-v4-pro)
  --early-stop, -s    Stop after N gens without improvement (default: 5)
  --checkpoint, -c    Continue from a checkpoint .py file
  --no-checkpoints    Don't save checkpoint files
  --quiet, -q         Suppress progress output
  --save-best         Save best program to a specific path