Purity in the EVM
This document gives what we mean by purity suitable for a signature validation contract. It gives resources for designing an on-chain purity-checking contract.
This document is not official advice. Errors may be present.
This document is available as a Git repository at .
Background
This document is the result of "reverse engineering" the following two contracts and the majority any credit attributed to this document is deserving of their authors:
- in by Vitalik Buterin.
- of the above Serpent Purity Checker by .
Vitalik's contract was ported to LLL by Alex with the intention that it would be employed to verify the "purity" of signature validation (valsig) contracts in the (now deprecated) EIP-1011 proposal. Before we go through the concept of purity, first we should read the purpose of valsig contracts.
Valsig contracts were to be employed to abstract the signature validation of vote and exit messages -- allowing validators to put in place arbitrarily intricate signature schemes rather of just relying upon transaction signatures (ECDSA).
Unfortunately, the use of arbitrary external valsig contracts opens the possibility for a damaging attack vector whereby a validator can, because they control signature validation, double-vote then prevent punishment by ensuring that the slash validation fails whilst the vote validation succeeds. Such an attack can be eliminated by reducing the space of possible valsig contracts to only those which are "pure": those which will always return the result given the same signature to validate.
Hence, EIP-1011 needed that each valsig contract must have its purity approved by an on-chain smart contract before it is permitted to be applied for validation. The contract that scanned and established valsig contact purity was called the purity checker and it is the focus of this document.
Definition of Purity
This contract offers the following definition of purity for an Ethereum smart-contract:
A contract is considered pure if it will always return the same result given sufficient gas for execution and the same transaction data and value. Namely, it may read the data and value fields of a transaction but no other transaction information, it may not read block information and it must not read from or write to storage.
There is some room for subjectivity in the definition of purity, in particular in what can be considered "inputs" to the "function" that is the smart contract. This definition covers only transaction data and value (, in the Yellow Paper) but given no concrete definition of what is transaction "context" (as opposed to transaction "input") what we mean by purity can be conceived which permits the origin address (derived from , and ) to be read as well. Such a definition is not compatible with signature validation contracts and is hence excluded from this document.
Detecting Impurity On-Chain
This document assumes that detecting the purity of a contract is going to be ran on the contracts bytecode. This is more accurate than performing the action on source code because it eliminates any quirks which may be rolled out during compilation. This also has the benefit of allowing on-chain verification of one contract by another as a contract may go and retrieve the bytecode of another contract and iterate through it inside the EVM.
The process of determining the purity of some bytecode will in general involve starting at the first byte (which must be an opcode), attempting to match it against the opcodes defined in the table, performing some action depending on the purity of that opcode (e.g., permit, deny or attempt address detection) and then repeating the process on the next opcode.
It helps to note that not each byte in some bytecode must be an opcode,
in place of that it may be a parameter supplied to a PUSH opcode. The Serpent contract
offered in the Background section gives an example of how
one can keep track of opcodes and parameters throughout the iteration process
to allow for back-searching of opcodes and parameters as needed for
address detection in call-type opcodes (see Address Detection
Techniques).
The rest of the document focuses on defining the purity categories for each opcode, outlining techniques that can be employed to deal with call-type opcodes and then offers some detail as to why certain opcodes have been categorised as pure or potentially-impure.
Impurity Categories
There are three classifications for impure opcodes: always impure, potentially impure call-type and future impure opcodes. Each category is described below.
Always Impure
These opcodes have no use other than to mutate state, return mutable state or give context about the execution environment. Any contract which contains an "always impure" opcode should be immediately considered impure.
Future Impure Opcodes
These opcodes are assumed to be reserved for future impure opcodes. At the time of writing, there is no formal declaration that this is the case and this judgement is solely based off the authors informal conversations with the Ethereum community.
Potentially Impure Call-Type
Call-type opcodes (see the table for a listing) may execute code at some other address. It is possible for an external call to be either pure or impure, depending on the address specified for the call. The use of a call-type opcode can only be considered pure if the address specified is:
- An address that has already been worked out to be pure.
- Any of the precompile addresses within the range of
0x0000000000000000000000000000000000000001to0x0000000000000000000000000000000000000008. Note: the purity of these contracts is yet to be confirmed.
See the Address Detection Techniques section for some techniques for extracting the address supplied to a call-type opcode from bytecode.
Any call to an externally-owned (non-contract) address should be considered impure. This is because it can potentially have impure code deployed to it.
Impure Opcode Table
| Opcode Value | Mnemonic | Impurity Category |
|---|---|---|
0x31 | BALANCE | Always Impure |
0x32 | ORIGIN | Always Impure |
0x33 | CALLER | Always Impure |
0x3a | GASPRICE | Always Impure |
0x3b | EXTCODESIZE | Always Impure |
0x3c | EXTCODECOPY | Always Impure |
0x40 | BLOCKHASH | Always Impure |
0x41 | COINBASE | Always Impure |
0x42 | TIMESTAMP | Always Impure |
0x43 | NUMBER | Always Impure |
0x44 | DIFFICULTY | Always Impure |
0x45 | GASLIMIT | Always Impure |
0x46 - 0x4F | Range of future impure opcodes | Future Impure Opcodes |
0x54 | SLOAD | Always Impure |
0x55 | SSTORE | Always Impure |
0xf0 | CREATE | Always Impure |
0xff | SELFDESTRUCT | Always Impure |
0xf1 | CALL | Potentially Impure Call-Type |
0xf2 | CALLCODE | Potentially Impure Call-Type |
0xf4 | DELEGATECALL | Potentially Impure Call-Type |
0xfa | STATICCALL | Potentially Impure Call-Type |
* 0xfb | CREATE2 | Always Impure |
* Opcodes which were not implemented at the time of writing, but the author has an expectation they will be implemented in the future.
Address Detection Techniques
Call-type opcodes (see the table for a listing) can only be considered pure if they call a particular set of addresses (see Potentially Impure Call-Types). So, to permit some call-type opcodes it is needed to establish the called address from the bytecode. This section describes methods which may be employed to find the address supplied to the call-type opcode with certainty.
The code which may place an address on the stack for call-type opcode can be arbitrarily involved and only discoverable by executing said code. To allow purity checking within a single Ethereum transaction the techniques here are simplistic and will give false positives (indicating impurity). Even so, these techniques should never produce false negatives (indicating purity).
Techniques are gave in a Python-like pseudo-code and concrete examples can be found in the two contracts specified in the Background section.
Convenience Functions
First two convenience functions are declared; get_opcode(n) and
get_last_opcode_param(n).
Convenience Function get_opcode(n)
Returns the n'th opcode declared in the subject bytecode[].
If n is out of bounds of bytecode[] the function returns None.
Example:
ADD = 0x01
PUSH2 = 0x61
bytecode = [PUSH2, 2, 1, ADD]
get_opcode(0)
# 3
get_opcode(2)
# NoneConvenience Function get_last_opcode_param(n)
Returns the final parameter supplied to the n'th opcode declared in
the subject bytecode[].
If n is out of bounds of bytecode[] or the n'th opcode does not have
parameters the function returns None.
Example:
ADD = 0x01
PUSH2 = 0x61
bytecode = [PUSH2, 2, 1, ADD]
get_last_opcode_param(0)
# 1
get_last_opcode_param(1)
# None
get_last_opcode_param(2)
# NoneAddress Detection Functions
Four functions are now declared which return an address if a given pattern
of opcodes is found to precede a call-type opcode. If all of these functions
return None, then the contract should be assumed to be impure.
Each function takes an input c which is the index of the call-type opcode in
question. It is assumed that the on-chain purity checking contract is iterating
over the bytecode in question and each time it detects a call-type opcode it
runs these functions to attempt to detect the address being called.
Address Detection Function #1
PUSH1 = 0x60
PUSH32 = 0x7f
def address_detector_1(c):
if PUSH1 <= get_opcode(c-2) <= PUSH32:
return get_last_opcode_param(c-2)
else:
return NoneAddress Detection Function #2
SUB = 0x03
GAS = 0x5a
PUSH1 = 0x60
PUSH32 = 0x7f
def address_detector_2(c):
if (get_opcode(c-1) == SUB and
get_opcode(c-2) == GAS and
PUSH1 <= get_opcode(c-3) <= PUSH32):
return get_last_opcode_param(c-3)
else:
return NoneAddress Detection Function #3
GAS = 0x5a
SWAP1 = 0x90
def address_detector_3(c):
if (get_opcode(c-1) == GAS OR
get_opcode(c-1) == SWAP1):
return get_last_opcode_param(c-2)
else:
return NoneAddress Detection Function #4
DUP1 = 0x80
DUP16 = 0x8f
def address_detector_4(c):
if (DUP1 <= get_opcode(c-1) <= DUP16):
return get_last_opcode_param(c-2)
else:
return NoneOpcode Listing
This section contains an opcode-by-opcode listing of each defined opcode. For each opcode the following is gave:
- Summary: a brief account of what the opcode does.
- Impurity Reasoning: a reference demonstrating impurity reasoning.
- Possible Attack: a scenario which assumes some attacker has deployed a contract and wishes to be able to have some pre-determined or ad hoc control of the return result of the contract. This section does not exhaustively list possible attacks, it simply offers an example for demonstrative purposes.
Specifications of opcodes can be found in Appendix H of the Ethereum Yellow Paper.
BALANCE
Summary: Returns the balance of some address. References: Impurity Reasoning: reads state. Possible Attack: An attacker may influence the return value of a contract call by altering the balance of some external account.
ORIGIN
Summary: Returns the address of the sender of the transaction which
triggered execution. In Solidity, this is tx.origin.
References:
Impurity Reasoning: reads illegal transaction context.
Plausible Attack: An attacker may influence the return value of a contract
call by varying the private key with which a transaction is signed.
CALLER
Summary: Returns the address immediately responsible for the execution. In
Solidity, this is msg.sender.
References:
Impurity Reasoning: reads illegal transaction context.
Possible Attack: An attacker may influence the return value of a contract
call by varying the private key with which a transaction is signed or with an
intermediary contract to alter the CALLER value.
GASPRICE
Summary: Returns the current gas price. References: Impurity Reasoning: reads illegal transaction context. Possible Attack: An attacker may influence the return value of a contract call by via some means to alter the gas price (e.g., immediately controlling block proposers).
EXTCODESIZE
Summary: Returns the size of the code held at some address. References: Impurity Reasoning: reads state. Plausible Attack: An attacker may influence the return value of a contract. call by deploying code to some pre-computed address.
EXTCODECOPY
Summary: Copies some amount of code at some address to some position in memory. References: Impurity Reasoning: reads state. Possible Attack: An attacker may influence the return value of a contract call by deploying code to some pre-computed address.
BLOCKHASH
Summary: Returns the hash of some past block (within the prior 256 complete blocks). References: Impurity Reasoning: reads state. Possible Attack: An attacker may influence the return value of a contract call by controlling some portion of block proposers and selecting block hashes based upon how they will influence the contract call.
COINBASE
Summary: Returns the beneficiary address of the block. References: Impurity Reasoning: reads state. Plausible Attack: An attacker may influence the return value of a contract call by controlling some portion of block proposers and declaring the beneficiary address based upon how it will influence the contract call.
TIMESTAMP
Summary: Returns the timestamp of the block. References: Impurity Reasoning: reads state. Plausible Attack: An attacker may influence the return value of a contract call by controlling some portion of block proposers and declaring the timestamp based upon how it will influence the contract call.
NUMBER
Summary: Returns the number of the block (count of blocks in the chain since genesis). References: Impurity Reasoning: reads state. Plausible Attack: An attacker may influence the return value of a contract call by selecting in which block a transaction should be contained.
DIFFICULTY
Summary: Returns the block difficulty. References: Impurity Reasoning: reads state. Plausible Attack: An attacker may influence the return value of a contract call by assuming some control of the collective hash rate and modifying it based upon how it will influence the contract call.
GASLIMIT
Summary: Returns the block gas limit. References Impurity Reasoning: reads state. Possible Attack: An attacker may influence the return value of a contract call by with some means to alter the gas limit (e.g., immediately controlling block proposers or spamming the network).
SLOAD
Summary: Returns a word from storage.
References
Impurity Reasoning: reads state.
Possible Attack: At the time of writing the author is not aware of any
attack with SLOAD if all other purity directives are followed. Even so,
attacks could be imagined if combined with the SLOAD opcodes (other
attacks may be possible).
SSTORE
Summary: Saves some word to storage.
References:
Impurity Reasoning: reads and mutates state.
Possible Attack: At the time of writing the author is not aware of any
attack with SSTORE if all other purity directives are followed. Even so,
attacks could be imagined if combined with the SSTORE or GAS opcodes (other
attacks may be possible).
CREATE
Summary: Creates a new account given some code.
References:
Impurity Reasoning: reads and mutates state.
Plausible Attack: At the time of writing the author is not aware of any
attack via CREATE if all other purity directives are followed. Still,
attacks could be imagined if combined with the EXTCODESIZE opcode (other
attacks may be possible).
SELFDESTRUCT
Summary: Registers the account for deletion, sending remaining Ether to some address. References: Impurity Reasoning: reads and mutates state. Possible Attack: An attacker may self-destruct a contract, causing all future calls to it to fail.
CALL
Summary: Message-calls to some address. References: Possible Impurity Reasoning: Executes code from another account. Possible Attack: An attacker may call an impure contract and use its return data.
CALLCODE
Summary: Execute the code of some other account via the state of this account. References: Possible Impurity Reasoning: Executes code from another account. Possible Attack: An attacker may callcode an impure contract and read or mutate state.
DELEGATECALL
Summary: Execute the code of some other account with the state of this
account whilst retaining the same values for sender and value.
References:
Possible Impurity Reasoning: Executes code from another account.
Plausible Attack: An attacker may delegate an impure contract and read or
mutate state.
STATICCALL
Summary: Message-calls to some address without persisting state modifications. References: Possible Impurity Reasoning: Executes code from another account. Plausible Attack: An attacker may call an impure contract and use its return data.
CREATE2
This opcode has not been implemented at the time of writing.
Summary: Creates a new account given some code and some nonce (as opposed
to CREATE which uses the current account nonce).
References: .
Impurity Reasoning: Reads and mutates state.
Plausible Attack: An attacker could craft a contract which succeeds the
first time it is called, but fails all other times.