Guide

Calculate a SegWit txid and wtxid with Python

Parse one synthetic SegWit transaction, remove its witness serialization and calculate both displayed identifiers offline with Python.

9 min readTransactions
Calculate a SegWit txid and wtxid with Python

Two identifiers from one transaction

A Segregated Witness transaction can have two identifiers. The txid hashes the traditional serialization without the SegWit marker, flag or witness fields. The wtxid hashes the complete witness serialization. They refer to the same transaction, but they commit to different bytes.

This guide parses a small synthetic transaction and calculates both identifiers with Python's standard library. The fixture is deliberately not a valid spend. It contains a made-up previous output and arbitrary witness items, so it must never be broadcast. That makes it useful for examining serialization without a wallet, private key, node, network connection or funds.

BIP 141 defines the two serializations and their identifiers. Bitcoin Core 31.1's transaction code serializes witness data only when requested and present. Those sources and the Bitcoin Core 31.1 release were checked on 9 October 2026. Version 31.1 was the current stable release when checked.

Prerequisites and expected result

You need Python 3 and a text editor. No package installation or internet connection is required. Create a disposable directory, save the program below as segwit_ids.py, review it, then run:

python3 segwit_ids.py

The tested environment used Python 3.12.14. A successful run reports a 69-byte total serialization, a 61-byte base serialization and two different identifiers:

txid=116102982e4852a9fe78c24f4c62e7b584378b956c83303d0d51b7b3b0fe6fb1
wtxid=8d195fe5fde6b872a64a027a8e33902a67cbdd802585e7a00d9dab19a072622a

It also rejects three malformed inputs: a truncated transaction, a wrong marker and a trailing byte.

The complete program

#!/usr/bin/env python3

from hashlib import sha256

RAW_HEX = (
    "02000000"
    "0001"
    "01"
    + "11" * 32
    + "00000000"
    "00"
    "ffffffff"
    "01"
    "e803000000000000"
    "01"
    "51"
    "02"
    "01"
    "01"
    "02"
    "0203"
    "00000000"
)


class Reader:
    def __init__(self, data):
        self.data = data
        self.offset = 0

    def take(self, size):
        end = self.offset + size
        if end > len(self.data):
            raise ValueError("truncated transaction")
        part = self.data[self.offset:end]
        self.offset = end
        return part

    def varint(self):
        prefix = self.take(1)[0]
        if prefix < 0xFD:
            return prefix, bytes([prefix])
        widths = {0xFD: 2, 0xFE: 4, 0xFF: 8}
        body = self.take(widths[prefix])
        value = int.from_bytes(body, "little")
        minimum = {0xFD: 0xFD, 0xFE: 0x10000, 0xFF: 0x100000000}[prefix]
        if value < minimum:
            raise ValueError("non-canonical compact size")
        return value, bytes([prefix]) + body


def copy_varbytes(reader):
    size, encoded_size = reader.varint()
    return encoded_size + reader.take(size)


def strip_witness(raw):
    reader = Reader(raw)
    base = bytearray(reader.take(4))
    if reader.take(2) != b"\x00\x01":
        raise ValueError("expected SegWit marker 00 and flag 01")

    input_count, encoded = reader.varint()
    if input_count == 0:
        raise ValueError("transaction has no inputs")
    base.extend(encoded)
    for _ in range(input_count):
        base.extend(reader.take(36))
        base.extend(copy_varbytes(reader))
        base.extend(reader.take(4))

    output_count, encoded = reader.varint()
    base.extend(encoded)
    for _ in range(output_count):
        base.extend(reader.take(8))
        base.extend(copy_varbytes(reader))

    for _ in range(input_count):
        item_count, _ = reader.varint()
        for _ in range(item_count):
            copy_varbytes(reader)

    base.extend(reader.take(4))
    if reader.offset != len(raw):
        raise ValueError("unexpected bytes after locktime")
    return bytes(base)


def displayed_hash(payload):
    return sha256(sha256(payload).digest()).digest()[::-1].hex()


if __name__ == "__main__":
    raw = bytes.fromhex(RAW_HEX)
    base = strip_witness(raw)
    txid = displayed_hash(base)
    wtxid = displayed_hash(raw)

    assert len(raw) == 69
    assert len(base) == 61
    assert txid != wtxid

    print(f"total_bytes={len(raw)}")
    print(f"base_bytes={len(base)}")
    print(f"base_hex={base.hex()}")
    print(f"txid={txid}")
    print(f"wtxid={wtxid}")

    invalid = {
        "truncated": raw[:-1],
        "wrong marker": raw[:4] + b"\x01" + raw[5:],
        "trailing byte": raw + b"\x00",
    }
    for label, candidate in invalid.items():
        try:
            strip_witness(candidate)
        except ValueError as error:
            print(f"PASS rejected {label}: {error}")
        else:
            raise AssertionError(f"FAIL accepted {label}")

Read the fixture in wire order

The first four bytes, 02000000, encode transaction version 2 in little-endian order. The next two bytes are the SegWit marker and flag, 00 01. They announce the extended serialization. They are included in the wtxid calculation but omitted from the txid calculation.

The next 01 says there is one input. Its previous transaction hash consists of 32 repeated 11 bytes. The following four bytes select output index zero. A compact-size value of zero means the scriptSig is empty, and ffffffff is the sequence. These values are only structural test data. They do not identify a spendable coin.

The next 01 says there is one output. e803000000000000 is 1,000 satoshis encoded as an unsigned 64-bit little-endian value. The one-byte output script is 51, the opcode OP_TRUE. Again, this is a parser fixture, not a payment.

The witness begins with 02, meaning two stack items. The first item is one byte, 01. The second is two bytes, 0203. The final four zero bytes are the locktime.

Build the base serialization safely

The program does not remove byte patterns with a text replacement. That would be unsafe because 0001 or witness-looking bytes can occur elsewhere. Instead, Reader advances through the transaction according to its field boundaries.

Compact-size integers determine the input count, output count, script lengths, witness item counts and witness item lengths. varint also rejects longer encodings of small values. copy_varbytes keeps each length prefix together with its payload when the field belongs in the base serialization.

strip_witness copies the version, inputs, outputs and locktime. It reads but does not copy the marker, flag or witness. It rejects a zero input count, truncated fields and any bytes after locktime. The resulting 61 bytes are:

020000000111111111111111111111111111111111111111111111111111111111111111110000000000ffffffff01e803000000000000015100000000

Hash and display the identifiers

Bitcoin uses SHA-256 twice for both identifiers. Hash bytes are conventionally displayed in the reverse order from the 32-byte digest returned by the hashing function. That is why displayed_hash reverses the digest before converting it to hexadecimal.

The txid passes the 61-byte base serialization to displayed_hash. The wtxid passes all 69 bytes. The eight-byte difference consists of the two-byte marker and flag plus six bytes of witness serialization. Changing only a witness item changes the wtxid while leaving the txid unchanged.

The result was independently recalculated with Node.js 24.19.0 using its built-in crypto module and an explicitly assembled base serialization. Both implementations produced the same two identifiers and byte counts.

Check the full tested output

total_bytes=69
base_bytes=61
base_hex=020000000111111111111111111111111111111111111111111111111111111111111111110000000000ffffffff01e803000000000000015100000000
txid=116102982e4852a9fe78c24f4c62e7b584378b956c83303d0d51b7b3b0fe6fb1
wtxid=8d195fe5fde6b872a64a027a8e33902a67cbdd802585e7a00d9dab19a072622a
PASS rejected truncated: truncated transaction
PASS rejected wrong marker: expected SegWit marker 00 and flag 01
PASS rejected trailing byte: unexpected bytes after locktime

The Python source SHA-256 was 8edaaa4bfd72dea2be37a98cceeadc2ce4fd74f02b38fc591cb50de053570e08. The complete output SHA-256 was ee3d0f4e394b0a7aa634935f9eab38d26720f326ab556737c492210faaa7b190. These hashes identify the exact local fixture and output that were tested. They do not make copied code trustworthy.

Troubleshooting

The identifiers are reversed: apply byte reversal only after the second SHA-256. Reversing the transaction bytes, or reversing after each hash, calculates something else.

The txid equals the wtxid: that is normal for a transaction without witness data. For this fixture, equality means the complete serialization was probably hashed twice or the marker, flag and witness were not removed from the txid input.

The parser reports truncation: check every compact-size length and the final locktime. A single missing byte can make a later length consume the wrong field.

A real transaction uses marker 00 with another flag: this teaching parser accepts only the currently defined 01 flag. Use maintained Bitcoin software for general network data rather than broadening the example by guesswork.

An explorer shows only one identifier: txid remains the common transaction reference used by block explorers and wallet interfaces. The wtxid is important for witness-aware relay and block commitments, but many interfaces do not display it.

Limits that matter

Calculating an identifier does not validate scripts, signatures, amounts, consensus rules or whether referenced outputs exist. This fixture is intentionally not a valid spend. Do not broadcast it and do not adapt it by inserting a real previous output or witness.

The program handles the exact current SegWit marker and flag and parses compact-size lengths, but it is an educational boundary check, not a replacement for Bitcoin Core's transaction decoder. BIP 141 also gives the coinbase transaction a special all-zero wtxid when constructing the witness commitment. Hashing a coinbase serialization with this function does not implement that separate rule.

Use the exercise to see which bytes each identifier commits to. For production software, use a maintained Bitcoin library, official test vectors and review that covers parsing, consensus validation and error handling together.

Newsletter

Bitcoin, without the noise

What happened in Bitcoin, what it actually changes, and the sources so you can check us. One issue at a time, straight to your inbox.

  • One email per issue, never a drip campaign
  • No tracking pixels and no shared addresses
  • Unsubscribe from any issue in one click

Get the next issue

One email per issue, no tracking pixels, and unsubscribe from any of them. We do not share your address. Privacy policy