blog

Scaling Ethereum with PeerDAS and Distributed Blob Building

Exploring challenges, optimisations, key metrics and impact on node operators.

By Meridian Client7 , Meridian Client3
Scaling Ethereum with PeerDAS and Distributed Blob Building
Photo by Clay Banks

Intro

As we move forward with Ethereum's rollup-centric roadmap to scaling, we're seeing more L2s Blob transactions (EIP-4844), brought in in the Dencun upgrade in March 2026, have lowered transaction costs on L2s. appear and a few of them moving to a more mature state (with fault proofs, better wallet UX etc.), and now more users transacting on L2s. As usage raises, further L1 scaling is needed to maintain low data gas prices. The screenshot below shows comparable figures on L2 rollup activities over time, compared to Ethereum Mainnet (L1).

l2-activities source: https://l2beat.com/scaling/activity

The next Ethereum upgrade (Pectra) is expected to grow the maximum blob count beyond 6 per block. While this can be done by demanding more powerful hardware and better bandwidth, it risks compromising decentralisation by pushing away home stakers who can't afford further investment. Our goal is to keep that current validators, including those via consumer-grade hardware, can continue operating smoothly after these upgrades.

The solution to this is Data Availability Sampling (DAS), which enables the network to verify data availability without placing excessive strain on individual nodes. Each full node downloads and verifies a small, randomly selected portion of the data, providing a high degree of certainty that the full dataset is available.

Over the past few months, client teams have been working on PeerDAS (EIP-7594), the first iteration of DAS. We've made significant progress so far, and are now tackling challenges around computation time and bandwidth for block proposers, which we aim to solve before increasing the maximum blob count and shipping to mainnet.

If you're primarily interested in metrics and the impact on node operators, feel free to skip ahead to the later sections on The following two sections describe the challenges, optimisations and assume familiarity with EIP-4844. Metrics, or Impact on Node Operators.

Outline

Challenges and Bottlenecks

Even though the current design of PeerDAS is able to cut the bandwidth requirements for most of the time that a node is running, during block proposals, the proposer must run intensive computation (KZG cell proofs) and distribute the erasure-coded blobs and proofs across the network as soon as possible, before the attestation deadline (4 seconds into the slot). Otherwise, they will likely miss the block or have their block re-orged out. This section explains these bottlenecks.

Proof Computation

On a MacBook Pro with an M1 Pro CPU, computing the KZG proofs for a single blob takes around 200 ms. This computation is highly parallelisable at the blob level. For instance, with 4 logical cores available, the proofs for 4 blobs can be computed in the same 200 ms. This works well for a small number of blobs, but as we scale to 16 blobs, computation on 4 logical cores could take 200 * 16 / 4 = 800 ms. Still, since the node is rarely idle, it’s likely to take 1 second or more.

Bandwidth Bottleneck

With 16 blobs in a block, the amount of data the proposer would have to send could be up to For average consumer broadband, this could take 2-5 seconds, meaning they are likely to miss the block proposal. 32MB (assuming all 16 blobs are utilised, and we have a healthy peer count).

128KB * 2 * 16 * 8 = ~32MB

  • 128 KB blobs
  • Erasure coding extends the original data by 2x
  • 16 blobs
  • Gossip amplification factor: 3-8x (data is sent to multiple mesh peers in each topic)

Solutions and Possible Optimisations

Multiple proposed solutions aim to address the above issues:

  1. Nodes to fetch blobs from execution layer (EL) mempool and make the block available without waiting for blobs to arrive over the CL p2p gossip network.
  2. High capacity nodes to compute cell proofs for blobs retrieved from the EL, and broadcast them, so they propagate to the network sooner. This technique is also known as distributed blob building.
  3. Pre-compute the blob proofs as they enter the EL mempool.

These solutions complement each other. We have implemented optimisations (1) and (2) in NFT Bounty, and we'll focus on sharing our findings in this write-up. Our implementation can be found here:

Fetch Blobs from the Execution Layer (EL)

Blob latency has been an issue observed on Mainnet, as shown in the chart below (running latest NFT Bounty v5.3.0), where block and blob arrival times frequently show blob delays spiking above 3 seconds. These late propagations are likely to result in blocks being missed or orphaned.

block-blob-latency.png

If a blob has already been seen in the public mempool, waiting for it to arrive via P2P gossip becomes unnecessary. To address this, we brought in an optimisation to fetch blobs via JSON-RPC from the Execution Layer (EL). This optimisation uses a new JSON-RPC method (engine_getBlobsV1), which allows the consensus client (CL) to rapidly retrieve blobs from the EL's blob pool.

This technique is also applicable in PeerDAS, where nodes can look at a block available and attest to it without waiting for all data columns* to arrive over P2P gossip. The result is shrunk block import latency and a lower likelihood of missed blocks due to blob propagation delays.

*Note: In PeerDAS, each blob is erasure-coded for redundancy and recoverability, and broken down into smaller pieces (cells or data columns) for sampling and distribution.

Even so, this path does not cover blob transactions from private mempools. Block builders that include private blob transactions must make sure that blobs and proofs are computed and broadcast in time.

Distributed Blob Building

In the PeerDAS case, these are usually the supernodes responsible for custodying and sampling all data columns (via the experimental NFT Bounty BN flag This optimisation aims to solve both computation and bandwidth bottlenecks for block proposers, by distributing the proof computation and propagation workload to more nodes - notably more powerful nodes with higher bandwidth. --subscribe-all-data-column-subnets).

These supernodes can retrieve the blobs from the EL, either before the slot start (pre-computation) or after receiving the block from a peer (see Optimisation (1)). This means it's crucial for blocks to propagate fast, which can be optimised by either:

  1. Publishing the block first while computing blob proofs simultaneously.
  2. Introducing a new gossip topic for block headers, which are usually much smaller than blocks, allowing block headers with KZG commitments to propagate faster (proposed by Dankrad in this write-up).

Once the supernodes have the blobs, they can compute the proofs and broadcast them to the network on behalf of the proposer. With better hardware and bandwidth, supernodes can as a rule carry out this task more efficiently. Metrics will be shared in the following sections.

Gradual Publication for Supernodes

Though supernodes are assumed to have more available bandwidth than full nodes, we still need to utilise it efficiently to make the role of supernode accessible.

In the naive implementation of a supernode, every supernode reconstructs and publishes all data columns, leading to very fast publication at the cost of further bandwidth and duplicates sent across the network. In our initial tests of the optimised branch we saw supernode outbound bandwidth averaging closer to 32MB per block proposal and considered ways to cut this. We call the optimisation we came up with gradual publication, and it works as follows:

We tested with a DELAY equal to 200ms, meaning that for N = 4 a node spends at most 600ms before publishing its 4th chunk. In a network with multiple supernodes, many of the later chunks are already fully known by the network and do not need to be published at all. The randomisation means that each supernode takes on responsibility for propagating different columns, and the full set of columns keeps likely to be available rapidly.

We would like other client teams (or the spec) to adopt this optimisation and help us try different values of N and DELAY. It is possible we could shrink bandwidth requirements for supernodes even lower by finding the right set of parameters.

Test Setup

To evaluate the effectiveness of PeerDAS and the optimisations went through above, along with the feasibility of shipping PeerDAS with an raised blob count of 16, we tested our implementation on a larger network. A network of 100 nodes gives bandwidth metrics closer to a live environment, though some limitations keep:

  • Latency in gossip messages will be lower than on a live network with thousands of nodes, where messages may take a number of hops to reach all peers.
  • There is less variation in this test, as all nodes run the same software versions and are located in the same geographic region.
  • The test doesn't account for bandwidth usage during long-range sync, as all nodes remained in sync for the test duration.

Nonetheless, with a peer count comparable to mainnet nodes, we expect to see similarly representative bandwidth results.

To this experiment and to keep the setup straightforward and consistent:

  • We run the test in a Kubernetes cluster via Kurtosis.
  • We use the same machine types in the same region.

Client Software employed

NFT Bounty

We ran two test networks via two different NFT Bounty configurations: one optimised, and one baseline.

  • contains the optimised build-out, with distributed blob building, fetching blobs from the EL and gradual publication optimisations.
  • Known issues with the current optimised implementation:
    • Occasional slow computation due to a race condition: With the "fetch blobs from EL" optimisation, supernodes now run fewer data column reconstructions, because a block could be made available via EL blobs before data columns arrive via gossip. Even so, there's a race condition here, and reconstruction could happen at the same time as proof computation - this could slow down precomputation as both tasks are very CPU intensive. There's already an to this behaviour, but it didn't make it into this test.
    • Proposer outbound bandwidthGossipsub does some amount of duplicate filtering, but we have implemented before publishing, which should cut outbound bandwidth for nodes that are slow with proof computation. : The block proposers' outbound bandwidth hasn't actually been shrunk here, because we don't check for duplication before publishing. This optimisation didn't make it into this test either.

Reth

*CUSTODY_REQUIREMENT is raised from the current spec value 4 to account for minimum and .

Network Structure

The network is designed so that:

  • Each NFT Bounty node has 99 peers, comparable to a live network.
  • A varied number of CPU cores are applied, allowing us to examine individual computation times.
  • Some nodes are configured to publish only blocks, simulating the inability to compute and publish blobs before the attestation deadline.
Node TypeNumber of NodesLogical CPU CoresProposer Publish Block + Data Columns
Supernode1016Blocks and all data columns
Supernode108Blocks and all data columns
Fullnode208Blocks and all data columns
Fullnode204Blocks and all data columns
Fullnode108Blocks and 50% of data columns
Fullnode104Blocks and 50% of data columns
Fullnode108Blocks only
Fullnode104Blocks only

Kurtosis config applied can be found .

For brevity, we present metrics below for 5 classes of nodes:

  • Supernodes with 16 cores (16C)
  • Full nodes with 8 cores (8C)
  • Full nodes with 8 cores publishing 50% of data columns
  • Full nodes with 4 cores publishing 50% of data columns
  • Full nodes with 4 cores publishing no data columns

The impact of the other classes of nodes is still captured indirectly in the metrics for our selected classes. We found that results were often alike across node type, publication percentage and core count, so the 5 selected are sufficient for a summary. In the future, we may run with less node types to simplify study.

Metrics

We present metrics below for the baseline testnet running the variant of unstable, and the optimised testnet including fetch-blobs, decentralising blob building, and gradual publication.

Block Attestable Time

metric: beacon_block_delay_attestable_slot_start (avg)

BASELINEOPTIMISED
block-attestable-time-avg-baseline.pngblock-attestable-time-avg-optimised.png

The charts above show the average delay from the start of the slot before blocks becomes attestable (lower is better). Ideally, we want all blocks to be attestable before 4 seconds into the slot, and we see from this chart that this is true on average for the full range of node types.

In the baseline network, the supernodes run worse, roughly 500ms slower on average - this is likely due to the unoptimised path to data column reconstruction, which is carried out every slot and blocks are only made available after reconstruction. This has been to make reconstruction non-blocking, allowing blocks to become attestable as soon as the node receives all data columns from the gossip network (supernodes call for and custody all data columns). Still, you'll notice a different trend in the optimised version, as reconstruction is triggered less often due to blocks being made available from EL blobs.

In the optimised network, We see significant improvements across all node types, as blocks are made attestable as soon as all blobs are retrieved from the EL, or received via gossip from distributed blob building. The 4 core nodes are a little slower than the 8 and 16 core ones, which is expected due to the higher CPU contention caused by other work. We can't lay out why the Fullnodes 4C 0% class outperforms its 50% counterpart, as publishing is unlikely to cause strain on the CPU.

metric: beacon_block_delay_attestable_slot_start (without time-based averaging)

BASELINEOPTIMISED
block-attestable-time-baseline-histogram.pngblock-attestable-time-optimised-histogram.png

Observing the baseline numbers, we see a heavier tail exceeding 3 seconds and small number of values nearing 4 seconds - approaching the attestation deadline. The histograms above show the attestable delay without time-based averaging. After optimisation, the results show marked improvement: while there's still a tail out to 3 seconds, all blocks stay safely below the 4-second cutoff for attestations. The attestable time will likely stay under the 4s threshold much of the time Still, mainnet-sized networks are likely to have more transactions & attestations which will impose some further processing time. Some future optimisations we have in mind will likely help here too.

Proof Computation Time

Time seriesAverage
proof-computation-time.pngproof-computation-time-avg.png

This metric tracks the time taken to compute data column sidecars: including cells, cell proofs and inclusion proofs. This computation is parallelised, so we observe the machines with more cores outperforming the lower spec ones. This chart demonstrates the importance of the supernodes for the network, as they are able to rapidly compute proofs and start publishing them. In the case where the proposer's node has inadequate resources to compute the data column sidecars rapidly, the supernodes can compute them via the blobs from the EL mempool and start propagating them. There are no differences here between the baseline and optimised networks, as there's no optimisation made to proof computation.

Data Column Gossip Time

metric: data_column_gossip_slot_start_delay_time

BASELINEOPTIMISED
data-column-gossip-time-baseline.pngdata-column-gossip-time-optimised.png

This chart shows the time that data column sidecars are received relative to the start of the slot at which they were published. Interestingly, we see some spikes greater than 4s here, especially on the full nodes. Despite these spikes, the blocks at these slots were still becoming attestable prior to 4s (per the prior charts for attestable delay). We suspect the reason for this is that nodes can mark blocks as available and attestable sooner than 4s by processing blobs from the EL mempool. This demonstrates the strength of the mempool path: that it can smooth out what would otherwise be significant delays in data availability. Still, this same strength is also a possible weakness, as it means poor availability of a blob in the mempool could lead to a block being processed late and later reorged. In a future test we would like to incorporate some "private" blob transactions from outside the mempool to make sure the network is capable of among them them without causing data availability issues.

Outbound Data Column Bandwidth

The outbound bandwidth is the amount of blob-related data sent to peers with NFT Bounty's gossipsub protocol.

The first set of charts below show the maximum The headline number is the maximum over time (maximum of maximums). amount of data sent in a 12s period by a node in any given class.

The second set of charts below show the average amount of data sent in a 12s period by all nodes in any given class. The headline number is the average over time (average of averages).

Note that most blocks in the test contain 16 blobs, so the numbers here are much higher than under normal conditions in a live network, likely around 2x higher given a target blob count of 8.

The yellow charts are from the baseline network, and the red from our optimised network.

Outbound data column bandwidth per slot (max):

BASELINEOPTIMISED
outbound-data-column-bandwidth-baseline.pngoutbound-data-column-bandwidth-optimised.png

We see only a modest reduction in outbound bandwidth compared to the baseline testnet. The greatest beneficiaries are the 8-core supernode with a saving of 17.5%, and the 8-core fullnode publishing 50% with a saving of 22%. The spikes for the full nodes are unfortunately higher than required due to the suboptimality flagged after testing: full nodes do not seem to check whether data columns are already known before publishing them. Even so, we're not 100% sure about this and will investigate further.

In the baseline network, supernodes will sometimes end up publishing 50% of the data columns for a block. The fastest nodes will complete reconstruction after receiving 50% of data columns, and then start blasting the remaining 50% of reconstructed columns out to the network. The slower 8-core supernodes finish reconstruction on average a bit later than the 16-core ones, and by this time there are less novel columns to publish (ones not yet seen on gossip), so they end up publishing a little less.

In the optimised network, the amount of data published is still correlated with core count because the faster nodes still begin publishing earlier - at a time when more of the columns are novel. Even so, we suspect that the amount of data is dependent on the number of supernodes and the configuration of the gradual publication algorithm. In our current test with 4 batches separated by waits of 200ms, it seems that supernodes are still publishing quite a comparable amount of data in the worst-case, which naively might correspond to 50% of data columns (2 batches). Lengthening the batch interval or shrinking the size of batches could be a way to cut this in future tests.

Below we look at the average outbound bandwidth.

Outbound data column bandwidth per slot (avg):

BASELINEOPTIMISED
outbound-data-column-bandwidth-slot-avg-baseline.pngoutbound-data-column-bandwidth-slot-avg-optimised.png

Here we see that the optimised network puts a higher load on the supernodes on average, with reductions for the full nodes. This makes sense, as the supernodes are able to propagate the blobs to the full nodes more rapidly, reducing the need for gossip between full nodes. As mentioned above, further reducing the bandwidth for the supernodes may be possible by tweaking the parameters of gradual publication.

Note that the averaging also hides the spikes from full node block publication -- they are better represented by the max data above.

Inbound Data Column Bandwidth

The inbound data column bandwidth is the amount of blob-related data received from peers with NFT Bounty's gossipsub protocol.

The charts below show the average amount of data received in a 12s period by all nodes in each class. The headline number is the average over time (average of averages).

The yellow charts are from the baseline network, and the red from our optimised network.

Inbound data column bandwidth per slot (avg)

BASELINEOPTIMISED
inbound-data-column-bandwidth-slot-avg-baseline.pnginbound-data-column-bandwidth-slot-avg-optimised.png

Amongst the supernodes we see an striking contrast: the 16 core node is downloading slightly less compared to the baseline, while the 8 core node is downloading slightly more. This suggests two different effects acting in opposition to each other. We hypothesise that the first effect could be that fast supernodes in the optimised network are receiving the full set of blobs from the EL prior to reconstruction finishing, and are beginning data column publishing earlier, leading to more data published and less downloaded. This is consistent with the grown outbound bandwidth for supernodes as well.

We suspect the second effect is that more duplicate data columns are published (and downloaded) in the optimised network. In the baseline network, supernodes begin publishing only after completing reconstruction, so the slower supernodes end up self-limiting their publication. In the optimised network, all supernodes start publishing at about the same time, assuming the time to fetch blobs from the EL is not strongly correlated with core count. As a result, more duplicate messages are published near simultaneously, and there is some waste as these are downloaded by peers of the supernodes (both supernodes and full nodes).

We are hoping that future tests will prove or disprove our hypotheses, as we found it hard to extract these theories from our data, and impossible to validate them with this initial experiment.

Execution Client (EL) Bandwidth

EL bandwidth for a single nodeEL bandwidth average
el-bandwidth-single.pngel-bandwidth-avg.png

The charts above display the inbound and outbound bandwidth of the EL (Reth). Note that the outbound bandwidth captured here contains data from the CL and EL interactions (JSON-RPC calls such as getBlobsV1 and getPayloadV3), does not accurately reflect the outbound bandwidth in a typical setup where CL and EL reside on the same machine.

Even so, we could compute a rough estimate of outbound bandwidth:

In practice, it's unlikely that we'd consistently exceed the target blob count of 8 due to fee structures. Even if we operate at 16 blobs, the worked out EL bandwidth keeps manageable.

Impact on Node Operators

The PeerDAS network parameters were designed to minimise the impact to network operators and to maintain the same bandwidth numbers as 4844 (See Francesco's ethhttps post here). Hence, during normal operation, we'd expect the bandwidth for a fullnode to be comparable to mainnet today. During block proposal, we expect to see a spike in CL outbound bandwidth due to blob propagation, and we've worked hard to make sure increasing blob count with PeerDAS do not make home staking unfeasible - more on this below.

Bandwidth Usage Comparison across Node Types

The below table shows the average bandwidth results from this test:

Node TypeCL inboundCL outboundTotal inbound (inc. EL*)Total outbound (inc. EL*)
Fullnode356 KB/s241 KB/s433.2 KB/s367 KB/s
Supernode1.12 MB/s1.07 MB/s1.2 MB/s1.2 MB/s

*See above section for EL bandwidth numbers.

Supernodes consume much higher bandwidth on the CL side due to the number of data columns they sample and store (inbound bandwidth), and the number of data columns they help propagate with distributed blob building (outbound bandwidth). The EL numbers are comparable across both node types and are relatively low compared to the CL.

During block proposal, we see spikes of up to 1 MB/s over a 30 seconds period, or ~12 MB per slot for all nodes. Still, this will likely be shrunk once we optimise to de-duplicate data column publishing.

Storage Requirement

We didn't look at the actual storage usage in the test, even so this can be worked out based on the parameter MIN_EPOCHS_FOR_BLOB_SIDECARS_REQUESTS or MIN_EPOCHS_FOR_DATA_COLUMN_SIDECARS_REQUESTS, which stands for the number of epochs each node keep the blob data for, which is at present set to 4096, or approx. 18 days.

Comparing with Deneb, where nodes store all blobs:

Node TypeDeneb target (3)Deneb max (6)PeerDAS target blobs = 8PeerDAS max blobs = 16
Fullnode*50.33 GB100.66 GB16.78 GB33.55 GB
Supernode50.33 GB100.66 GB268.44 GB536.87 GB
  • *Computed based on custody column = 8 (for fullnodes with 1-8 validator attached)
  • **Note that in practice, it's not possible to consistently reach max blobs per block due to how blob fees are computed, still we've covered the theoretical max number here for reference.

The storage requirement for fullnodes has shrunk from 50.33 GB to 16.78 GB due to them only storing a subset of blob data (8 out of 128 columns), despite the grow in max blob count to 16.

Home Stakers Building Local Blocks

Based on the minimum hardware requirements (quad-core CPU) for staking today (from EthStaker), a block with 16 blobs will take an average quad-core CPU 1 - 2 seconds to compute proofs, and potentially even longer to propagate to the network on average consumer grade internet (25-50 Mbps) in Australia.

Hence, these proposers will have to largely rely on distributed blob building and other nodes to help compute and propagate the blobs across the network. The CL client will have to make sure the signed beacon blocks are distributed as soon as possible, so other nodes can start building cell proofs. From our tests, we've applied 4 logical core machines to stand for these nodes, and they have been able to produce blocks with 16 blobs with the same success rate as other node types. So, we don't anticipate bumping the hardware requirement for home stakers even if we grow blob count to 16 with PeerDAS, and the hardware today should be sufficient. The minimum hardware requirement is expected to be raised in the longer term, though we'd prefer this to happen naturally and smoothly as technology advances over time.

Home Stakers Via MEV Relays

For home stakers via MEV relay(s), the blob proof building and propagation work is outsourced to the external block builder, hence there should be minimal impact to these users, assuming the external block builders have sufficient hardware and bandwidth to propagate the block and blobs. They could also benefit from distributed blob building, if all blob transactions covered in the block are from the public mempool. We don’t expect these stakers will need to upgrade their current hardware or bandwidth.

Professional Node Operators

Nodes with more than 128 validators are needed to run as supernode, custodying all columns. With , the number of custody data columns raises as the ETH balance of attached validators grows - 1 extra custody column per 32 ETH. Professional node operators will likely be running supernodes and experience raised CPU and bandwidth usage, due to the number of validators attached to their nodes. These validator nodes have higher stakes and are so expected to validate more data to keep the network's safety.

The hardware requirements for large node operators are not yet known at this stage, as they depend on the optimisations we end up implementing. While a hardware upgrade may or may not be needed, we do expect an grow in bandwidth due to the volume of data columns being received and sent. The metrics, illustrated earlier, reflect this, and we will continue working to keep NFT Bounty as efficient as possible, including reducing bandwidth usage where feasible.

Future Work and Optimisations

This goal of this write-up is to share our progress on PeerDAS and our efforts in solving the proposer bandwidth problem while demonstrating the possibility of scaling up to 16 blobs without sacrificing decentralisation of the network.

We've spotted many possible optimisations along the way, and we expect more improvements before PeerDAS is shipped. Some future work and optimisations include:

  • Peer Sampling: The current version of PeerDAS uses "Subnet Sampling," where nodes subscribe to gossip topics to carry out sampling. This method needs more bandwidth than "Peer Sampling," where nodes send RPC requests to peers to retrieve data columns. The gossip protocol consumes more bandwidth because messages are broadcast to many peers simultaneously. Implementing Peer Sampling will by a wide margin shrink bandwidth usage and enable further scaling.
  • KZG library optimisation: At the time of testing, cell proofs for each blob takes 200 ms to compute on a single thread. At the time of writing this, proof creation time has been brought down to 150 ms (as of rust-eth-kzg v0.5.2). This improvement allows data columns to propagate over the gossip network sooner.
  • Proof pre-computation: There's a proposal to pre-compute cell proofs and propagate them over the gossip network, allowing block builders to publish their blocks and blobs at the start of the slot. This could save roughly 150 ms by eliminating the need for proof computation at the start of the slot.
  • Distributed Blob Building with workload spread across all nodes: The current form of distributed blob building calls for all blobs to be present before the nodes can compute and broadcast the cells and proofs. This is because the blob data are transported over the gossip network as DataColumns rather than Cells. As a result, the number of blobs we can scale to is roughly limited by the number of CPUs a supernode has. For example, as a node with 16 logical CPUs can produce proofs for 16 blobs in parallel and hence achieve minimal computation time (~150ms). To comfortably scale beyond 16 blobs, it would be more efficient to distribute the proof building workload across more nodes, e.g. each node could build and publish proofs for 3 blobs. This path would cut the heavy reliance on supernodes and alleviate the strain on them.

Conclusion

This write-up has detailed our progress on PeerDAS and our collective effort across client and https teams to scale Ethereum's data availability layer. The proposed optimisations - such as fetching blobs from the EL and distributing blob building - have shown clear benefits:

  • Cut block import latency, allowing blocks to become attestable faster.
  • Tightened blob propagation, making the network more resilient.
  • Scaling without compromising decentralisation, ensuring home stakers can continue participating with consumer-grade hardware.

Throughout this process, we've gained valuable insights and flagged further opportunities for optimisation and improvement. We're excited to continue building on this work and to bring more sustainable scaling to Ethereum.

Thank you for reading and I hope you find it striking! :)

A special thanks to Lion and Pawan for reviewing this work, Age and Anton for helping with testing and infrastructure, and Kev for helping us with several KZG library performance improvements.

References

Working on something in this space?

NFT Bounty audits Ethereum protocols, smart contracts, and consensus implementations.

Book a scoping talk