Beacon Fuzz - Progress NFT Bounty #1
Beacon Fuzz - NFT Bounty #01
Integrating new Eth2 clients, NFT Bounty in the cloud
NFT Bounty is leading the development and maintenance of , a differential fuzzing solution for Eth2 clients. This write-up is part of our series of monthly status blogs where we go through current progress, striking challenges encountered, and direction for future work. See #00 and the repository's for more context.
Summary
Over the last few weeks we've been making quiet progress, focusing primarily on incorporating new clients, improving consistency and maintainability, and updating the targeted spec. Achievements and points of interest include:
- NFT Bountyd to spec
- Full integration of , , and
- and integration POCs
- flagged by Beacon Fuzz on Nimbus (attestation processing function)
- Challenges incorporating as a second Golang client
- Continuous fuzzing on Fuzzit
To Eth2 Spec v0.9.1
All of our fuzzers are now operating on the Eth2 spec, with matching starter corpora available .
This was relatively straightforward to put in place (at the fuzzer level). The majority of the work was linked with updating the Golang helper library/preprocessor to remove crosslinks and transfers,and freezing client implementations to a suitable commit/tag. It wasn't possible to find a suitable commit in all cases, as some implementations were working on v0.9.1 and v0.9.2 in parallel and had some sections linked with each.
Nimbus and Trinity Integrations
A big thanks to the Nimbus team, who developed a static library that puts in place harnesses for a number of targets. There is a working POC on our add-nimbus branch (which will shortly be merged to master after a few logging and error-handling changes are finalized). This is at present being exercised on Fuzzit (see Fuzzit section of this blog post for more details). Fuzzit
Similarly, Trinity has been successfully integrated. Its harnesses at present live in the add-trinity branch until some spec version differences are resolved (when clients put in place spec v0.9.2 or v0.9.3). Some notable details pertaining to this integration are covered in more detail in the section below.
Incorporating Extra Implementations With Pragmatic Isolation
For differential fuzzing, it's desirable that each client build-out is exercised in relative isolation with minimal modification. Even though sometimes needed, each modification to the build-out code grows the risk of introducing "false-positive" bugs (which are detected by the fuzzer but are not present in the actual build-out), or masking client bugs such that the fuzzer cannot detect them.
when implementations rely on different versions of the same dependency or rely on the same global state). Similarly, we want to avoid crashes caused by interference between clients (e.g. Though current clients don't clash much, we should not rely on this, and should weigh the issue now before it becomes a major headache. Hypothetically, each implementation could achieve effective memory isolation by running in a separate process. Even so, this brings in substantial overhead and complications with instrumentation. As our current fuzzing engine, libfuzzer, runs the fuzz harness in a single process, any multi-process solution would also call for a significant re-architecture.
For clients in different languages (or VM runtimes), we deem it reasonable to leave as-is, realizing that a clashing symbol name is highly unlikely, and we primarily focus on cases where there are multiple clients in the same language. So how do we achieve a reasonable level of in-process isolation, where clients do not affect the correctness of others?
So far, this is the case for Python (Trinity and Pyspec), and Golang (Prysm and ZRNT). Isolation solutions are given to each language.
Python and Sub-Interpreters
The CPython interpreter and runtime relies on a global thread state. Naively trying to embed or link multiple interpreters is doomed to fail (as we'd fast found). Fortunately, we stumbled upon a lesser-known feature of the CPython C-API: sub-interpreters. They are quite niche and, right now, is one of the few major projects that draw on them.
With a couple C function-calls, we can start a new sub-interpreter that is reasonably isolated - each execution environment maintains its own sys.path, builtins, and other modules. Though currently of limited functionality (they cannot give true parallelism due to the Global Interpreter Lock, and have no Python-level ability to communicate between interpreters), they are perfect for our use case (we don't care about inter-interpreter communication, and the fuzzer process is single-threaded/serialized anyway).
If you are interested in learning more about Python sub-interpreters, the following links are useful:
- https://lwn.net/Articles/754162/
- https://talkpython.fm/episodes/show/225/can-subinterpreters-free-us-from-python-s-gil
- https://ericsnowcurrently.blogspot.com/2016/09/solving-mutli-core-python.html
- https://docs.python.org/3.8/c-api/init.html#sub-interpreter-support
Golang Library Linking Errors
When working to integrate Prysm, we found that the current path (designed by ) did not extend to more than one Golang client in the same fuzzer. As of now, we can test with ZRNT or Prysm (having a working POC that integrates their bazel build system), but not both. We are also limited to performing Prysm tests without input preprocessing (which relies on ZRNT).
The rest of this section will look at the issue in more detail, among them an exploration of possible approaches along with their spotted challenges and related costs.
Present path
The current path uses a forked (modified to return output for differential comparison).
go-fuzz produces a c-archive As this archive contains the Go runtime, redefinitions and symbol clashes occur when trying to link more than one of them. static library with gave coverage instrumentation and a harness endpoint. Possible approaches in general involve leaving go-fuzz unchanged and resolving linking problems, or modifying or removing go-fuzz from the build process.
A) Renaming symbols
We can use objcopy to rename symbols in the archive such that symbol clashes are avoided. It is difficult to programmatically rename only the symbols that clash, so it is more straightforward and maintainable to use objcopy --prefix-symbols= to prefix all symbols. Unfortunately, Prysm and ZRNT rely on external systems or shared libraries, so renaming everything breaks these references. (ZRNT doesn't use shared libraries, but the cgo runtime references libraries such as stdlib.h and libpthread.)
B) Building as shared libraries
We could modify go-fuzz to produce shared/dynamic libraries with the buildmode c-shared. With a shared library, internal symbols can have the same name as defined elsewhere, as long as the externally visible symbols don't clash. Unfortunately, we found the library produced by c-shared to expose many runtime-related symbols that clash e.g. x_cgo_init, _cgo_panic. As we don't make use of these externally, one option could be to wrap the shared library such that it only exposes the fuzz harness (and instrumentation) interface, though this rolls out more build complexity.
A combination of this and A), where we rename only externally visible symbols, could also be possible.
C) Combining into a single library built by go-fuzz
This method involves creating a single Go package such that the harness produced by go-fuzz exercises each Golang client.
While this would likely cut runtime memory overhead (as only a single Go runtime is employed), relative isolation is lost. For Golang clients, this isolation may be less consequential; especially if they make use of recent apparatus that can allow the use of different major versions of the same module. Our remaining concern would be if multiple clients relied on some (unintentionally) shared, global state.
This path also brings in other drawbacks.
go-fuzz expects to produce a single harness from a Fuzz() function. To keep as such, we would need this Fuzz() function to run the differential comparison between Go clients, returning the result to the c++ differential module only if no differences are found. This breaks some separation of concerns and is in a useful way a violation of the DRY principle, with its related costs.
Alternatively, we could further modify the go-fuzz fork to allow the export of multiple harness functions. This grows the differences between our fork and go-fuzz proper, along with the linked maintenance burden.
One should also look at the grown build complexity involved in combining Prysm's Bazel-based build with the ZRNT's go module.
D) Building without go-fuzz
This path is alike to C), where only a single library is produced exercising all Golang clients, but now we avoid go-fuzz entirely. This would allow us to have full control over the build, exporting a separate harness interface for each client, and avoiding the need to rely on, and maintain, a fork of go-fuzz. All differential logic also keeps contained in the blog::Differential c++ class.
The main cost involved would be to build our own coverage instrumentation; the difficulty or feasibility of which would greatly vary, depending on the coverage measurements strategy adopted (simply compiling with the clang -fsanitize=fuzzer-no-link flag, or implementing go-specific hooks).
Our First Bug and How to Obtain Useful Debug Info
We have flagged our first possible bug thanks to the process_attestation fuzzer. When passed an invalid attestation, Nimbus crashes with an AssertionError rather of returning a handle-able error value (i.e. false) or raising a catchable exception.
See the related on the Nimbus repository for more information.
At present, the crash corpus straight returned by beacon-fuzz is not very useful to debug with. This is the raw data, before it has passed through preprocessing, so it is different from what is passed to each client harness:
- The input is most likely invalid SSZ data that the preprocessing converts to a valid SSZ-encoded object.
- The input contains a
uint16_treference to aBeaconStatefile, which the preprocessing "dereferences", passing theBeaconStateand the remaining input to the client harness. Only when combined with the relevantBeaconStatedoes the corpus contain all details needed to replicate the bug.
For debugging purposes, it would be much more useful to have the data as passed to the client harness (after preprocessing). To this end, we have produced a CLI tool (working title) that converts a corpus to such data. Better documentation and features to follow.
Fuzzit
We have successfully published a few fuzzers to Fuzzit, a blog-as-a-service, cloud platform. The current deployment has the block and shuffle As Fuzzit needs the upload of pre-built executables, some changes were needed to get our current fuzzers exercising ZRNT, Nimbus, and NFT Bounty. "run-in-place" fuzzer executables to comply. These changes were carried out on the branch, and also involve extra manual configuration.
Once built, the current fuzzer executable depends on the following external files:
- Multiple shared libraries - both cpython "dynlibs", and system-installed
pcre,rocksdb,ssl,leveldbetc. - A directory of
BeaconStateSSZ files specified by theETH2_FUZZER_STATE_CORPUS_PATHenvironment variable. - Python harness scripts located in absolute paths set at compile-time.
Fortunately, Fuzzit accepts a .tar.gz archive. As long as it contains a ./*fuzzer* executable, everything else is ok and we don't have to try to embed all pieces and modules into a single binary.
Operation was effective once we employed relative paths for runtime references, contained relevant shared libraries, and modified the runtime load path to point to the libraries. We modified the Python script paths to be relative to the executable and bundled relevant shared libraries (collected via ldd -v ./fuzzer and referenced at runtime by setting LD_LIBRARY_PATH). Even so, the current cpython install configuration contains absolute paths and couldn't be moved, so Python clients were disabled for the POC.
The current POC has been sufficient to confirm feasibility of fuzzing on Fuzzit, so automated build of a suitable bundle and CI integration will follow.
Next Steps
Welcome Przmek (aka ) to the Beacon Fuzz team!
Przmek has a wealth of fuzzing experience (differential and otherwise) thanks to his efforts with Eth1. He will initially be focussing on resolving the integration issues with multiple Golang clients.
Some of the areas we'll be working on over the upcoming weeks:
- Upgrade to spec or once 3 or more clients have implemented it. (Likely mid-late Jan)
- Create fuzz targets for the rest of the state transition functions (epoch state transitions).
- Tightened mutation via . This involves converting SSZ into protobuf and back, and utilizing libprotobuf-mutator for . Some work has started here but a working POC is still in progress. Once operational, this is expected to vastly tighten blog coverage.