- Rust 94.9%
- Python 5.1%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
Byte-for-byte-compatible Rust port of the Sing-song encoding by Panayotis Vryonis (vrypan), with SIMD encode/decode (AVX-512 VBMI, AVX2), SHAKE-256 variants, CLI, and parity tests against the bundled reference.py. Source: https://blog.vrypan.net/2026/08/19/260819-sing-song/ |
||
| src | ||
| tests | ||
| .gitignore | ||
| bench-results-rust.txt | ||
| Cargo.lock | ||
| Cargo.toml | ||
| README.md | ||
| reference.py | ||
babu — the Sing-song codec in Rust
A fast, byte-for-byte-compatible Rust port of the Sing-song encoding, plus the original Python reference implementation.
Sing-song is a reversible encoding of arbitrary byte strings as pronounceable consonant–vowel syllables. Every 6-bit chunk maps to one syllable (
b d f g j k l m n p r s t v w z×a i o u); hyphens are cosmetic; complete encodings are self-sizing and canonical; an optional variant XORs the input with a deterministic SHAKE-256 mask.— Panayotis Vryonis (vrypan), Sing-song
let enc = babu::encode(&[0x51, 0xac, 0x07, 0x59, 0xfc, 0x4d]);
assert_eq!(enc, "kalo-tadu-komu-tigi");
assert_eq!(babu::decode(&enc).unwrap(), &[0x51, 0xac, 0x07, 0x59, 0xfc, 0x4d]);
Repository layout
Cargo.toml crate manifest (single dependency: sha3)
src/lib.rs re-exports the codec
src/babu.rs codec core: encode / decode / variants (SHAKE-256 masks)
src/babu/simd.rs SIMD tiers (AVX-512 VBMI, AVX2) + runtime detection
src/bin/babu.rs CLI mirroring the Python one (exit code 2 on errors)
src/bin/bench.rs throughput + memory-ceiling benchmark
tests/parity.rs byte-for-byte parity vs Python goldens
tests/data/*.txt golden corpora generated by running reference.py
reference.py the normative Python reference (verbatim from the blog)
reference.py is reproduced unchanged from the blog post so that the two
implementations live side by side; the golden corpora in tests/data/ were
produced by running it.
Usage
Rust
cargo run --release --bin babu -- encode 51ac0759fc4d # → kalo-tadu-komu-tigi
cargo run --release --bin babu -- encode 51ac0759fc4d -v 5 # variant suffix appended
cargo run --release --bin babu -- decode kalo-tadu-komu-tigi # → 51ac0759fc4d
Python reference (identical interface)
python3 reference.py encode 51ac0759fc4d
python3 reference.py decode kalo-tadu-komu-tigi
The codec in one paragraph
Input bytes are read as a big-endian bit stream and split into 6-bit chunks,
most-significant bit first: bits 5..2 select one of 16 consonants, bits
1..0 one of 4 vowels. L bytes produce n = ceil(8L/6) syllables; the last
chunk is zero-padded (canonical padding that must be zero on decode). A
complete encoding is self-sizing — L = floor(6n/8) — so no length metadata is
needed, and leading zero bytes are preserved. Variant v (0–15) XORs the
input with SHAKE-256("sing-song/variant" ‖ v), rendered as a trailing
two-letter suffix (al am an ar il … ur); XOR is self-inverse, so decoding
XORs again with the same mask.
Performance
cargo run --release --bin bench measures encode, decode (SIMD and forced
scalar), the variant paths and a memcpy ceiling over a 32 B → 32 MiB sweep.
Selected results on a Zen 5 machine (AVX-512 VBMI tier; full output in
bench-results-rust.txt):
| input | encode | decode (SIMD) | decode (scalar) | SIMD gain |
|---|---|---|---|---|
| 32 B | 0.033 µs (0.97 GB/s) | 0.074 µs | 0.10 µs | 1.4× |
| 1 KiB | 0.10 µs (10 GB/s) | 0.32 µs (3.2 GB/s) | 3.2 µs | 9.8× |
| 16 KiB | 1.2 µs (14 GB/s) | 4.4 µs (3.7 GB/s) | 51 µs | 11.8× |
| 64 KiB | 4.6 µs (14 GB/s) | 16.7 µs (3.9 GB/s) | 208 µs | 12.5× |
| 1 MiB | 74 µs (14 GB/s) | 269 µs (3.9 GB/s) | 3.3 ms | 12.2× |
Variant paths add a SHAKE-256 mask derivation + XOR pass and are bounded by
the scalar SHAKE-256 throughput (~0.5 GB/s): encode_variant ≈ 0.49 GB/s and
decode_variant ≈ 0.45 GB/s at 1 MiB.
For reference, the Python implementation needs ~25–29 ms to encode/decode 16 KiB and ~0.35–0.44 s at 64 KiB, and its cost grows quadratically (one giant integer shifted per syllable). The Rust port is ~21 000× faster at 16 KiB and ~76 000× faster at 64 KiB on encode — the margin widens with input size.
Optimizations
- No big integers. 3 bytes = 24 bits = 4 syllables, so the bit transforms stay in machine registers (the Python reference shifts one giant integer per syllable — O(n²) in the payload).
- SIMD encode. 48-byte blocks (16 groups) become 160 bytes of
cvcv-tokens in one pass: VBMI byte permutes to repack the 24-bit groups,vpermbconsonant/vowel lookups, a full-width interleave, and hyphen insertion viapermutex2var+mask_set1. - SIMD decode. The character→6-bit-value map (~2.67 text chars per output
byte) runs as byte-permute table lookups: AVX-512 VBMI (
vpermb, 64-byte tables keyed byc & 0x3F) or AVX2 (pshufb+ nibble/bit-4 split tables), with a scalar fallback — all three verified to agree. - SIMD repacking. 6-bit syllable bytes → output bytes via 16-/32-bit lane
shifts plus one
vpermb(~10 vector ops per 24 bytes). - Canonical-grouped fast path. Standard Sing-song text puts a hyphen after every four letters, so a 40-character block always holds exactly 32 letters; AVX-512 validates that layout and decodes straight from the text with no hyphen-stripping copy — hyphenated decode is faster than decoding the same text after a cleanup pass.
- SIMD hyphen stripping. For arbitrary (non-canonical) hyphen placement the
cleanup pass uses
vpcompressbinstead of a scalar filter. - Memory. Tables are tiny and cache-resident; output buffers are pre-sized exactly and written sequentially; unaligned loads throughout.
Testing
cargo test --release
- Unit tests: spec vectors, canonicality rejects, suffix semantics, variant round-trips (all lengths 0–63 × all 16 variants), SIMD-vs-scalar agreement (including over random junk strings), SIMD encode vs scalar across the block boundary, and SIMD hyphen-strip agreement.
tests/parity.rs: byte-for-byte parity withreference.pyover a golden corpus (590 variant encodings, 480 grouped encodings, and rejected strings), all generated by running the Python reference.- The CLI mirrors the Python one, including exit code 2 on invalid input.
Attribution and license
Sing-song was designed and specified by Panayotis Vryonis (vrypan) in the
post Sing-song: a speakable encoding for long numbers and keys;
reference.py is reproduced verbatim from that post.
The Rust port (babu) is licensed under MIT OR Apache-2.0 (see
Cargo.toml). No license is asserted over reference.py, which remains the
author's work.