No description
  • Rust 94.9%
  • Python 5.1%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Chris Roscher 36c72a0311
Sing-song codec in Rust (babu) + Python reference
Byte-for-byte-compatible Rust port of the Sing-song encoding by Panayotis
Vryonis (vrypan), with SIMD encode/decode (AVX-512 VBMI, AVX2), SHAKE-256
variants, CLI, and parity tests against the bundled reference.py.

Source: https://blog.vrypan.net/2026/08/19/260819-sing-song/
2026-09-04 11:24:34 +02:00
src Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
tests Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
.gitignore Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
bench-results-rust.txt Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
Cargo.lock Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
Cargo.toml Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
README.md Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00
reference.py Sing-song codec in Rust (babu) + Python reference 2026-09-04 11:24:34 +02:00

babu — the Sing-song codec in Rust

A fast, byte-for-byte-compatible Rust port of the Sing-song encoding, plus the original Python reference implementation.

Sing-song is a reversible encoding of arbitrary byte strings as pronounceable consonant–vowel syllables. Every 6-bit chunk maps to one syllable (b d f g j k l m n p r s t v w z × a i o u); hyphens are cosmetic; complete encodings are self-sizing and canonical; an optional variant XORs the input with a deterministic SHAKE-256 mask.

— Panayotis Vryonis (vrypan), Sing-song

let enc = babu::encode(&[0x51, 0xac, 0x07, 0x59, 0xfc, 0x4d]);
assert_eq!(enc, "kalo-tadu-komu-tigi");
assert_eq!(babu::decode(&enc).unwrap(), &[0x51, 0xac, 0x07, 0x59, 0xfc, 0x4d]);

Repository layout

Cargo.toml            crate manifest (single dependency: sha3)
src/lib.rs            re-exports the codec
src/babu.rs           codec core: encode / decode / variants (SHAKE-256 masks)
src/babu/simd.rs      SIMD tiers (AVX-512 VBMI, AVX2) + runtime detection
src/bin/babu.rs       CLI mirroring the Python one (exit code 2 on errors)
src/bin/bench.rs      throughput + memory-ceiling benchmark
tests/parity.rs       byte-for-byte parity vs Python goldens
tests/data/*.txt      golden corpora generated by running reference.py
reference.py          the normative Python reference (verbatim from the blog)

reference.py is reproduced unchanged from the blog post so that the two implementations live side by side; the golden corpora in tests/data/ were produced by running it.

Usage

Rust

cargo run --release --bin babu -- encode 51ac0759fc4d          # → kalo-tadu-komu-tigi
cargo run --release --bin babu -- encode 51ac0759fc4d -v 5     # variant suffix appended
cargo run --release --bin babu -- decode kalo-tadu-komu-tigi   # → 51ac0759fc4d

Python reference (identical interface)

python3 reference.py encode 51ac0759fc4d
python3 reference.py decode kalo-tadu-komu-tigi

The codec in one paragraph

Input bytes are read as a big-endian bit stream and split into 6-bit chunks, most-significant bit first: bits 5..2 select one of 16 consonants, bits 1..0 one of 4 vowels. L bytes produce n = ceil(8L/6) syllables; the last chunk is zero-padded (canonical padding that must be zero on decode). A complete encoding is self-sizing — L = floor(6n/8) — so no length metadata is needed, and leading zero bytes are preserved. Variant v (0–15) XORs the input with SHAKE-256("sing-song/variant" ‖ v), rendered as a trailing two-letter suffix (al am an ar il … ur); XOR is self-inverse, so decoding XORs again with the same mask.

Performance

cargo run --release --bin bench measures encode, decode (SIMD and forced scalar), the variant paths and a memcpy ceiling over a 32 B → 32 MiB sweep. Selected results on a Zen 5 machine (AVX-512 VBMI tier; full output in bench-results-rust.txt):

input encode decode (SIMD) decode (scalar) SIMD gain
32 B 0.033 µs (0.97 GB/s) 0.074 µs 0.10 µs 1.4×
1 KiB 0.10 µs (10 GB/s) 0.32 µs (3.2 GB/s) 3.2 µs 9.8×
16 KiB 1.2 µs (14 GB/s) 4.4 µs (3.7 GB/s) 51 µs 11.8×
64 KiB 4.6 µs (14 GB/s) 16.7 µs (3.9 GB/s) 208 µs 12.5×
1 MiB 74 µs (14 GB/s) 269 µs (3.9 GB/s) 3.3 ms 12.2×

Variant paths add a SHAKE-256 mask derivation + XOR pass and are bounded by the scalar SHAKE-256 throughput (~0.5 GB/s): encode_variant ≈ 0.49 GB/s and decode_variant ≈ 0.45 GB/s at 1 MiB.

For reference, the Python implementation needs ~25–29 ms to encode/decode 16 KiB and ~0.35–0.44 s at 64 KiB, and its cost grows quadratically (one giant integer shifted per syllable). The Rust port is ~21 000× faster at 16 KiB and ~76 000× faster at 64 KiB on encode — the margin widens with input size.

Optimizations

  • No big integers. 3 bytes = 24 bits = 4 syllables, so the bit transforms stay in machine registers (the Python reference shifts one giant integer per syllable — O(n²) in the payload).
  • SIMD encode. 48-byte blocks (16 groups) become 160 bytes of cvcv- tokens in one pass: VBMI byte permutes to repack the 24-bit groups, vpermb consonant/vowel lookups, a full-width interleave, and hyphen insertion via permutex2var + mask_set1.
  • SIMD decode. The character→6-bit-value map (~2.67 text chars per output byte) runs as byte-permute table lookups: AVX-512 VBMI (vpermb, 64-byte tables keyed by c & 0x3F) or AVX2 (pshufb + nibble/bit-4 split tables), with a scalar fallback — all three verified to agree.
  • SIMD repacking. 6-bit syllable bytes → output bytes via 16-/32-bit lane shifts plus one vpermb (~10 vector ops per 24 bytes).
  • Canonical-grouped fast path. Standard Sing-song text puts a hyphen after every four letters, so a 40-character block always holds exactly 32 letters; AVX-512 validates that layout and decodes straight from the text with no hyphen-stripping copy — hyphenated decode is faster than decoding the same text after a cleanup pass.
  • SIMD hyphen stripping. For arbitrary (non-canonical) hyphen placement the cleanup pass uses vpcompressb instead of a scalar filter.
  • Memory. Tables are tiny and cache-resident; output buffers are pre-sized exactly and written sequentially; unaligned loads throughout.

Testing

cargo test --release
  • Unit tests: spec vectors, canonicality rejects, suffix semantics, variant round-trips (all lengths 0–63 × all 16 variants), SIMD-vs-scalar agreement (including over random junk strings), SIMD encode vs scalar across the block boundary, and SIMD hyphen-strip agreement.
  • tests/parity.rs: byte-for-byte parity with reference.py over a golden corpus (590 variant encodings, 480 grouped encodings, and rejected strings), all generated by running the Python reference.
  • The CLI mirrors the Python one, including exit code 2 on invalid input.

Attribution and license

Sing-song was designed and specified by Panayotis Vryonis (vrypan) in the post Sing-song: a speakable encoding for long numbers and keys; reference.py is reproduced verbatim from that post.

The Rust port (babu) is licensed under MIT OR Apache-2.0 (see Cargo.toml). No license is asserted over reference.py, which remains the author's work.