No description
- Python 100%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
| framework | ||
| problems | ||
| templates | ||
| .gitignore | ||
| evolve.py | ||
| README.md | ||
Evolvium
AI-driven code evolution — the AI iteratively mutates a Python function to maximize a fitness score, keeping improvements and discarding regressions.
How it works
flowchart LR
%% ── input files ──
PM["📄 problem.md<br/>task description"]
SOL["📄 solution.py<br/>current code"]
EPY["📄 evaluation.py<br/>scoring logic"]
%% ── core loop ──
EVAL["🔧 Evaluate<br/>run code, compute score"]
CMP{"Score improved?"}
LLM[("🧠 LLM<br/>mutate code")]
CKPT["📁 checkpoints/<br/>historical snapshots"]
%% ── decision block: SAVE wraps DISCARD ──
subgraph outcome["After comparison"]
direction TB
SAVE["💾 Save checkpoint<br/>Update solution.py"]
DISCARD["🗑️ Discard<br/>(revert to previous)"]
end
PM -->|"reads problem description"| LLM
SOL -->|"current code"| EVAL
EPY -->|"scoring logic"| EVAL
EVAL -->|"fitness score"| CMP
CMP -->|"yes – keep mutation"| SAVE
CMP -->|"no – revert"| DISCARD
SAVE -->|"best code so far"| LLM
DISCARD -->|"previous code\n(unchanged)"| LLM
LLM -->|"mutated code"| EVAL
SAVE -.->|"checkpoint snapshot"| CKPT
%% ── styles ──
classDef input fill:#dbeafe,stroke:#3b82f6,color:#1e3a5f
classDef compute fill:#fef3c7,stroke:#f59e0b,color:#78350f
classDef llm fill:#ede9fe,stroke:#8b5cf6,color:#4c1d95
classDef success fill:#d1fae5,stroke:#10b981,color:#064e3b
classDef discard fill:#fee2e2,stroke:#ef4444,color:#7f1d1d
classDef storage fill:#f3f4f6,stroke:#6b7280,color:#1f2937
class PM,SOL,EPY input
class EVAL,CMP compute
class LLM llm
class SAVE success
class DISCARD discard
class CKPT storage
- Read
solution.py— the current best code (inside# EVOLVE-BLOCK-START/ENDmarkers) - Evaluate — run
evaluation.pyto compute a fitness score - Mutate — send the problem description + current code + score to an LLM, asking it to improve the code
- Compare — if the mutated code scores higher, it becomes the new best; checkpoint saved,
solution.pyupdated - Repeat for N generations
If the initial solution.py is broken or missing, the LLM gets it anyway and can produce a valid first version from scratch.
Quick start
# Test the loop without an LLM (mock returns code unchanged)
python evolve.py --problem-dir sum_target --llm mock
# Run with a real LLM (requires pi)
python evolve.py --problem-dir sum_target --llm pi --generations 5
# Use a different model
python evolve.py --problem-dir tammes --llm pi --model deepseek/deepseek-v4-pro
Creating a new problem
Each problem needs exactly 3 files:
| File | Purpose |
|---|---|
problem.md |
Plain-English description of the optimization task. The LLM reads this. |
solution.py |
Python code with # EVOLVE-BLOCK-START / # EVOLVE-BLOCK-END markers. The LLM only mutates code between these markers. |
evaluation.py |
Two functions: get_score(result) -> float (higher = better) and evaluate(data) -> (metrics, outputs). The harness calls get_positions(n, starting_positions) from solution.py. |
Copy templates/PROBLEM_TEMPLATE.py as a starting point, or copy an existing problem directory:
cp -r problems/sum_target problems/my_problem
# Edit problem.md, evaluation.py, solution.py
python evolve.py --problem-dir my_problem --llm mock # test
python evolve.py --problem-dir my_problem --llm pi # run
LLM backends
| Backend | Flag | Notes |
|---|---|---|
| pi | --llm pi |
Calls pi -p --no-session --no-tools as a subprocess. Default model: deepseek/deepseek-v4-flash. Set EVOLVIUM_MODEL env var or use --model. |
| mock | --llm mock |
Returns the input program unchanged. For testing the evolution loop. |
| openai | --llm openai |
Uses the OpenAI API. Requires OPENAI_API_KEY. Set OPENAI_MODEL env var or use --model. |
CLI reference
python evolve.py --problem-dir <name> [options]
Options:
--problem-dir, -d Problem directory name or path (required)
--generations, -g Number of generations (default: 10)
--llm, -l LLM backend: pi, mock, openai, auto (default: auto)
--model Override model (e.g. deepseek/deepseek-v4-pro)
--early-stop, -s Stop after N gens without improvement (default: 5)
--checkpoint, -c Continue from a checkpoint .py file
--no-checkpoints Don't save checkpoint files
--quiet, -q Suppress progress output
--save-best Save best program to a specific path