FalkorDB 6.0 is the first release built on our new Rust engine. Across more than 300 benchmark queries, it executes 39% fewer CPU instructions than the C engine it replaces. It closes 226 bugs, including 43 crashes and 6 memory-safety defects. And your existing Cypher queries, APIs, and integrations work without a single change.
Try 6.0
docker pull falkordb/falkordb:latest
Recap: The Rust Re-write
In August we introduced FalkorDB’s new core engine, rebuilt from the ground up in Rust. The project took 17 months and produced 80,000 lines of Rust across 357 merged pull requests.
We chose Rust because it gives a database both safety and control. It enforces memory safety at compile time, which removes entire categories of bugs before the code ever runs, while keeping the low-level performance a graph engine needs. It also lets our team build and ship new features faster and with more confidence.
We followed one rule throughout: make it work, make it stable, then make it fast.
We kept the ideas that define FalkorDB, graphs stored as sparse matrices and traversals run as matrix multiplication, and rebuilt them in Rust. The new engine reads the same data files as the previous one, so existing data loads as is.
Every change had to pass about 1,585 openCypher TCK scenarios, 1,322 flow tests from the previous engine’s own suite, and a dedicated concurrency suite before anyone reviewed it. Our engineers owned every design decision and reviewed every change, and coding agents handled well-scoped implementation work inside those guardrails.
Performance work started only once the engine was correct. New techniques like columnar batch execution and string interning brought the Rust engine to parity with the previous engine, and ahead of it on many workloads.
The rewrite also introduced two new foundations: a columnar runtime that processes rows in batches, and multi-version concurrency control, so reads never block and writes stay consistent. Read the full story: Rewriting FalkorDB in Rust.
What has changed since our rust re-write announcement
New effects version
FalkorDB 6.0 introduces a new replication format, effects v3.
6.0 introduces effects v3, a new replication format designed to keep every replica exactly in sync with its primary, and to prove it.
In the previous format, the primary sent a compact record of each change and the replica filled in the rest itself. Entity IDs are a good example: rather than receiving them, the replica worked them out from the order in which new nodes and edges arrived. This is efficient, and it works as long as both sides always make the same calculation. V3 removes that assumption entirely.
The core change is that the replica now verifies instead of inferring. A few examples of what that looks like:
Every node, edge, label, and attribute arrives with both its ID and its name. The replica uses the ID the primary assigned and confirms it matches, rather than generating its own and expecting the two to line up.
When an entity is deleted, the record includes the labels it held. The replica needs that information to apply the delete correctly, and once the entity is gone there is no other way to recover it.
If the replica receives something it can’t account for, such as an unknown field, a count that doesn’t add up, or a payload that fails its checksum, it refuses the change on the spot and reports it.
Because this is a property of the format and not of any single engine, we are also incorporating v3 to the C engine as well. The goal is for both engines to behave identically.
That is what makes it a safer upgrade path. Once both engines speak the same verified format, a Rust replica can follow a C primary through a live cutover, and switch back if you need to reverse the upgrade. Every change is checked as it arrives. If the two engines ever disagree, you find out at that exact moment, with a clear refusal you can act on.
Less work per query than the C engine
Every performance number in this post comes from a single benchmark harness that lives in the FalkorDB repository, right next to the engine code. It follows one rule: a result only counts if you get the same number when you measure it again.
The workload. More than 300 Cypher queries run against a fixed graph of 10,000 people connected in a ring, plus smaller datasets for indexes, constraints, and multi-edges. The queries range from simple expressions and single clauses to aggregations, multi-step pipelines, cartesian products, bulk writes and deletes of different sizes, and a write against a large set of edges.
We also track which parts of the engine these queries reach. Coverage is currently around 70%, and we use it to spot the gaps and generate new queries with LLMs to fill them.
What we measure. For every query:
CPU instructions executed by the server, using perf on Linux and proc_pid_rusage on macOS. This is the most stable number and the one we rely on most to find and fix performance issues.
Memory allocated, read from the memory allocator’s own statistics.
Optionally, cache misses and branch mispredictions from the CPU’s hardware counters, to explain why a specific query is slow.
Wall-clock time is recorded too, but we don’t use it to accept or reject a change. It varies too much from run to run, and it improves on its own when instructions and allocations go down.
Side by side with the C engine. The same harness runs the production C engine, so every query reports a Rust-to-C ratio. A ratio below 1.0x means Rust does less work. A performance regression gets caught in code review, the same way the TCK catches a correctness regression.
What changed since the Rust engine became the default
Since August 2, about 30 of the roughly 95 changes merged to main have been measured performance improvements. Most followed the same pattern: the benchmark flagged a query that could be faster, profiling showed the engine doing work the query didn’t need, and the fix removed that work.
By early September, the full suite ran at 0.61x of the C engine’s instructions. That’s 39% less work overall.
Measured changes
| Area | Change | Before → After | vs. C engine |
|---|---|---|---|
| Columnar execution | Evaluate whole expression trees a column at a time (#2628) | Arithmetic filter 13.2M → 4.6M instr (2.8x); OR/CASE 2.3–2.6x | — |
| Keep aggregates columnar with computed keys / DISTINCT (#2693) | count(DISTINCT …) 36.2M → 7.6M instr (4.8x) | 1.93x → 0.40x | |
| Keep edge-property aggregates vectorized (#2416) | sum(r.k) 3,191 → 1,578 instr/edge (2.0x) | ~2x better than C | |
| Call constant functions once per batch, not per row (#2712) | split+trim+replace 3.8x fewer instr | 2.83x → 0.74x |
Show 18 more
| Area | Change | Before → After | vs. C engine |
|---|---|---|---|
| Answer n:Label from the label matrix as a column (#2494, #2685) | Allocations 976 KB → 12 KB per 10k-row scan | Instr 2.23x → 0.30x; alloc 2.29x → 0.004x | |
| Stop allocating a grouping key per row (#2708) | −7.6% instr on grouped pipelines | — | |
| Code Efficiency | Fix reversed traversal-chain ordering in the planner (#2491) | 34.4M → 0.44M instr (78x) | → 1.05x |
| Size the first scan batch to a downstream LIMIT (#2757) | scan LIMIT 1 3.6x, 1-hop LIMIT 10 3.0x | — | |
| Skip edge-id lookup when nothing reads the edge (#2428) | Traversal + count 1.99x | — | |
| Stream cartesian-product branches (#2365) | 4-way product: 2.5 min / 20.8 GB → 25 ms / 28 MB | C: 27 ms | |
| Count degree without materializing edge ids (#2572) | Up to −12% per degree call | — | |
| Write path | Paged edge-endpoint index; writes copy a page, not the index (#2691) | Write on 10M-edge graph 114x cheaper; bytes per CREATE flat (~51x less at 200k edges) | — |
| Remove the unused live-node matrix (#2678) | Create after 500k deletes 1.88 → 0.019 ms (~99x) | Level with C | |
| Snapshot deleted edges from the node's own adjacency (#2642) | DELETE … RETURN at 80k edges 1.39 → 0.10 ms (13x), now flat | Was 5–7x slower | |
| Reuse iterators when cascading deletes (#2674) | write 1m 56.8B → 17.4B instr (3.3x) | 2.58x → 0.79x; small writes → 0.25x | |
| Reclaim deleted ids by rank, removing an O(N²) walk (#2377, #2641) | Create-after-delete up to 2.2x cheaper per node | write 1m 4.29x → 2.58x | |
| Linear multi-edge demotion (#2431) | 229,534 → 17,058 instr/pair (13x) | — | |
| Delete in O(deleted), not O(graph) (#2378) | delete node −25%, write 1 −16% | — | |
| Ingest & fixed costs | Stop quoting the whole query in discarded parse errors (#2442) | 6.7 MB bulk batch parse 1,810 → 310 ms (5.8x) | — |
| Parse parameters straight into values (#2451) | Same batch 341 → 52 ms (6.5x) | 13.7x → 2.1x | |
| Write the telemetry stream the way C does (#2493) | Per-query floor cut by ~58k instr | 1.46x → 1.06x | |
| Keep single-match hash-join slots inline (#2690) | Fewer allocations on every hash join | hash join 0.66x → 0.56x |
| Area | Change | Before → After | vs. C engine |
|---|---|---|---|
| Columnar execution | Evaluate whole expression trees a column at a time (#2628) | Arithmetic filter 13.2M → 4.6M instr (2.8x); OR/CASE 2.3–2.6x | — |
| Keep aggregates columnar with computed keys / DISTINCT (#2693) | count(DISTINCT …) 36.2M → 7.6M instr (4.8x) | 1.93x → 0.40x | |
| Keep edge-property aggregates vectorized (#2416) | sum(r.k) 3,191 → 1,578 instr/edge (2.0x) | ~2x better than C | |
| Call constant functions once per batch, not per row (#2712) | split+trim+replace 3.8x fewer instr | 2.83x → 0.74x | |
| Answer n:Label from the label matrix as a column (#2494, #2685) | Allocations 976 KB → 12 KB per 10k-row scan | Instr 2.23x → 0.30x; alloc 2.29x → 0.004x | |
| Stop allocating a grouping key per row (#2708) | −7.6% instr on grouped pipelines | — |
| Area | Change | Before → After | vs. C engine |
|---|---|---|---|
| Code Efficiency | Fix reversed traversal-chain ordering in the planner (#2491) | 34.4M → 0.44M instr (78x) | → 1.05x |
| Size the first scan batch to a downstream LIMIT (#2757) | scan LIMIT 1 3.6x, 1-hop LIMIT 10 3.0x | — | |
| Skip edge-id lookup when nothing reads the edge (#2428) | Traversal + count 1.99x | — | |
| Stream cartesian-product branches (#2365) | 4-way product: 2.5 min / 20.8 GB → 25 ms / 28 MB | C: 27 ms | |
| Count degree without materializing edge ids (#2572) | Up to −12% per degree call | — |
| Area | Change | Before → After | vs. C engine |
|---|---|---|---|
| Write path | Paged edge-endpoint index; writes copy a page, not the index (#2691) | Write on 10M-edge graph 114x cheaper; bytes per CREATE flat (~51x less at 200k edges) | — |
| Remove the unused live-node matrix (#2678) | Create after 500k deletes 1.88 → 0.019 ms (~99x) | Level with C | |
| Snapshot deleted edges from the node's own adjacency (#2642) | DELETE … RETURN at 80k edges 1.39 → 0.10 ms (13x), now flat | Was 5–7x slower | |
| Reuse iterators when cascading deletes (#2674) | write 1m 56.8B → 17.4B instr (3.3x) | 2.58x → 0.79x; small writes → 0.25x | |
| Reclaim deleted ids by rank, removing an O(N²) walk (#2377, #2641) | Create-after-delete up to 2.2x cheaper per node | write 1m 4.29x → 2.58x | |
| Linear multi-edge demotion (#2431) | 229,534 → 17,058 instr/pair (13x) | — | |
| Delete in O(deleted), not O(graph) (#2378) | delete node −25%, write 1 −16% | — |
| Area | Change | Before → After | vs. C engine |
|---|---|---|---|
| Ingest & fixed costs | Stop quoting the whole query in discarded parse errors (#2442) | 6.7 MB bulk batch parse 1,810 → 310 ms (5.8x) | — |
| Parse parameters straight into values (#2451) | Same batch 341 → 52 ms (6.5x) | 13.7x → 2.1x | |
| Write the telemetry stream the way C does (#2493) | Per-query floor cut by ~58k instr | 1.46x → 1.06x | |
| Keep single-match hash-join slots inline (#2690) | Fewer allocations on every hash join | hash join 0.66x → 0.56x |
* A ratio under 1.0x means the Rust engine executes fewer CPU instructions than the C engine on that query.
Hundreds of bugs closed
6.0 resolves 226 issues reported against the C engine, some going back to August 2023. Many were closed as a direct result of the move to Rust. The rest surfaced during compatibility testing, when the new engine ran against the openCypher TCK and the full existing test suite.
The largest group, about 166 issues, covers query correctness: cases where a result didn’t match what was expected or where behavior didn’t fully follow the Cypher specification. Another 43 were stability issues, such as segmentation faults and out-of-memory terminations, and 6 were memory-safety issues, including a use-after-free in the query plan cache and an out-of-bounds read during bulk loading. These are the kinds of bugs Rust prevents at compile time, and a big part of why we made the move. The remaining fixes cover 4 hangs or timeouts and 6 performance issues.
Limitations
6.0 does not yet include a supported upgrade path from 4.x. The upgrade and downgrade path ships in 6.2. Until then, we recommend running 6.0 on new deployments or alongside your existing setup.
Availability
FalkorDB 6.0 is available now on Docker Hub:
Try 6.0
docker pull falkordb/falkordb:latest
We will roll 6.0 out to FalkorDB Cloud gradually as we collect feedback. Try it on your workloads and tell us what you find on GitHub or Discord. For a conversation about running it, contact us.
Authors
-
Avi Avni is Chief Architect at FalkorDB, specializing in graph database architectures for generative AI and retrieval-augmented generation workflows. He brings over 11 years of startup consulting experience, previously designing GraphMatrix—a full Cypher graph database—at Sela and leading RedisGraph from inception to enterprise readiness for Redis’ Tier 1 clients.
-