Skip to content

Track remaining funding-payment consistency issues after #962 #1044

Description

@tnull

PR #962 addresses the in-process races between wallet sync, funding classification, and payment graduation. Several persistence and identity questions remain deliberately unresolved.

Known remaining problems:

  • Funding classification writes the payment record and pending metadata separately. A failure or crash between them can leave a payment without its candidate history or txid mappings.
  • The failed classification package has already been dequeued. Repair may depend on a later LDK rebroadcast, which is not guaranteed for every persisted signing state or counterparty-broadcast transaction.
  • If the missing metadata belongs to an RBF funding payment, wallet sync may not map a non-first candidate back to the stable payment ID. It can consequently create a duplicate generic payment or retain incorrect candidate figures.
  • If wallet sync records an active RBF candidate before classification completes, its generic record may use the active candidate’s txid as its payment ID. Classification subsequently uses the first candidate’s txid and does not currently reconcile the two records.
  • It remains unclear how long mappings for earlier, conflicting, and graduated candidates must be retained.

Relevant discussions:

Follow-up questions/checklist:

  • Define the consistency guarantees required when only part of a funding classification can be persisted.
  • Determine how incomplete classifications are detected and recovered after both write failures and restarts.
  • Clarify whether and when LDK is guaranteed to redeliver a funding broadcast after classification fails.
  • Define how wallet sync resolves every funding candidate txid to the intended payment ID.
  • Define how sync-created generic records are reconciled with later funding classification.
  • Determine how long candidate and conflicting-txid mappings must remain available, including after graduation.
  • Define the expected behavior when the counterparty broadcasts during or after a partial local classification.
  • Account for Retry user-initiated splices across restarts and disconnects #930’s generated splice payment IDs and Stop keeping all DataStore entries in-memory, add pagination #1024’s payment-store caching changes.
  • Add regression coverage for partial writes, restart recovery, non-first candidate confirmation, and distinct first/active candidate IDs.

This issue intentionally does not prescribe a storage or recovery design.

Activity

jkczyz commented on Aug 12, 2026

@jkczyz
Contributor

🤖 A few things from the splice work (#930) and the analysis around it that aren't captured here yet.

1. LDK rebroadcasts cause problems of their own.

The second bullet mentions relying on a later LDK rebroadcast for repair. Rebroadcasts also cause some of the problems this issue tracks. With the pinned LDK rev (9174965), once a 0conf splice is locked — splice_locked goes out at zero confirmations because a splice inherits the channel's minimum_depth — LDK rebroadcasts the still-unconfirmed funding transaction on every monitor-update completion until it confirms, including after restarts. These rebroadcasts come typed TransactionType::Funding { channels } rather than InteractiveFunding, so they go through the generic classification path, with amounts taken from the on-chain wallet rather than from the splice contributions. On current main:

I hit all three on #930's branch; a regression test there (0conf splice, then payments to drive monitor updates) reproduces them. I'm splitting the fixes for the first two, plus the test, out of #930 into a separate PR. The third checklist item should probably also ask what a redelivered broadcast looks like, since it may come with a different type and amounts. It may also be worth asking upstream whether these rebroadcasts could keep the InteractiveFunding type — though TransactionType is documented as best-effort, so we'd still need to handle it.

2. bump_channel_funding_fee after ANTI_REORG_DELAY.

Agreed in #962 (comment) but not tracked anywhere: fail early when the splice already reached ANTI_REORG_DELAY confirmations, using ChannelDetails::splice_details once the rust-lightning#4687 backport lands.

3. remove_payment leaves the pending entry behind for good.

remove_payment deletes only the payment-store record. Graduation is the only path that removes a pending entry, and it declines when the record is gone, so the entry and its txid mappings outlive the deleted payment and keep resolving those txids to the deleted payment's id.

4. bump_fee_rbf doesn't take funding_payment_update_lock.

It reads and writes both stores without the lock, so it can interleave with a concurrent classification — the same problem the sync arms had before they held the lock from id lookup through their writes. Pre-existing; never came up on #962.

Suggested checklist additions:

  • Handle rebroadcasts of a locked-but-unconfirmed funding transaction without overwriting or duplicating the classified record.
  • Resolve candidate txids to the stable payment id in classification, not only in wallet sync.
  • Make bump_channel_funding_fee fail once a splice has ANTI_REORG_DELAY confirmations (needs the ChannelDetails::splice_details backport).
  • Remove or repair a payment's pending entry when the payment is deleted via remove_payment.
  • Take funding_payment_update_lock in bump_fee_rbf.
  • Regression test: rebroadcast of a locked-but-unconfirmed (0conf) funding transaction.

jkczyz commented on Aug 12, 2026

@jkczyz
Contributor

1. LDK rebroadcasts cause problems of their own.

The second bullet mentions relying on a later LDK rebroadcast for repair. Rebroadcasts also cause some of the problems this issue tracks. With the pinned LDK rev (9174965), once a 0conf splice is locked — splice_locked goes out at zero confirmations because a splice inherits the channel's minimum_depth — LDK rebroadcasts the still-unconfirmed funding transaction on every monitor-update completion until it confirms, including after restarts. These rebroadcasts come typed TransactionType::Funding { channels } rather than InteractiveFunding, so they go through the generic classification path, with amounts taken from the on-chain wallet rather than from the splice contributions. On current main:

I hit all three on #930's branch; a regression test there (0conf splice, then payments to drive monitor updates) reproduces them. I'm splitting the fixes for the first two, plus the test, out of #930 into a separate PR. The third checklist item should probably also ask what a redelivered broadcast looks like, since it may come with a different type and amounts. It may also be worth asking upstream whether these rebroadcasts could keep the InteractiveFunding type — though TransactionType is documented as best-effort, so we'd still need to handle it.

Opened #1049 for this.

tnull commented on Aug 17, 2026

@tnull
CollaboratorAuthor

One concrete case from #1024: WalletEvent::TxReplaced reads the authoritative PaymentDetails before refreshing its pending-store mirror. With a bounded payment cache, that read can now reach the KVStore and fail after BDK has already applied the wallet update. update_payment_store then consumes the event and skips the remainder of the batch, while those events may not be generated again.

This is the same recovery gap, but triggered by a read failure rather than a partial write. We should define how wallet-sync events are retained or retried after payment-store failures instead of falling back to potentially stale pending metadata. See the #1024 review thread (#1024 (comment)).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions