Repository navigation
Track remaining funding-payment consistency issues after #962 #1044
Description
Activity
🤖 A few things from the splice work (#930) and the analysis around it that aren't captured here yet.
1. LDK rebroadcasts cause problems of their own.
The second bullet mentions relying on a later LDK rebroadcast for repair. Rebroadcasts also cause some of the problems this issue tracks. With the pinned LDK rev (9174965), once a 0conf splice is locked — splice_locked goes out at zero confirmations because a splice inherits the channel's minimum_depth — LDK rebroadcasts the still-unconfirmed funding transaction on every monitor-update completion until it confirms, including after restarts. These rebroadcasts come typed TransactionType::Funding { channels } rather than InteractiveFunding, so they go through the generic classification path, with amounts taken from the on-chain wallet rather than from the splice contributions. On current main:
classify_fundingrecords every broadcast unconditionally, without checking for wallet activity. Both peers broadcast the fully-signed transaction, so the side that didn't contribute — whose wallet sees nothing — gets a zero-amount payment record for each rebroadcast.- If the rebroadcast txid matches the record's id (the first candidate), the reclassification has no guard against
Fundinglanding on anInteractiveFundingrecord: the wallet-derived amounts and the less specific type overwrite the negotiated ones. - If it doesn't (a non-first RBF candidate),
classify_fundingfalls back toPaymentId(txid)without consultingfind_payment_by_txidand creates a duplicate. So classification has the same txid-resolution problem the checklist raises for wallet sync.
I hit all three on #930's branch; a regression test there (0conf splice, then payments to drive monitor updates) reproduces them. I'm splitting the fixes for the first two, plus the test, out of #930 into a separate PR. The third checklist item should probably also ask what a redelivered broadcast looks like, since it may come with a different type and amounts. It may also be worth asking upstream whether these rebroadcasts could keep the InteractiveFunding type — though TransactionType is documented as best-effort, so we'd still need to handle it.
2. bump_channel_funding_fee after ANTI_REORG_DELAY.
Agreed in #962 (comment) but not tracked anywhere: fail early when the splice already reached ANTI_REORG_DELAY confirmations, using ChannelDetails::splice_details once the rust-lightning#4687 backport lands.
3. remove_payment leaves the pending entry behind for good.
remove_payment deletes only the payment-store record. Graduation is the only path that removes a pending entry, and it declines when the record is gone, so the entry and its txid mappings outlive the deleted payment and keep resolving those txids to the deleted payment's id.
4. bump_fee_rbf doesn't take funding_payment_update_lock.
It reads and writes both stores without the lock, so it can interleave with a concurrent classification — the same problem the sync arms had before they held the lock from id lookup through their writes. Pre-existing; never came up on #962.
Suggested checklist additions:
- Handle rebroadcasts of a locked-but-unconfirmed funding transaction without overwriting or duplicating the classified record.
- Resolve candidate txids to the stable payment id in classification, not only in wallet sync.
- Make
bump_channel_funding_feefail once a splice hasANTI_REORG_DELAYconfirmations (needs theChannelDetails::splice_detailsbackport). - Remove or repair a payment's pending entry when the payment is deleted via
remove_payment. - Take
funding_payment_update_lockinbump_fee_rbf. - Regression test: rebroadcast of a locked-but-unconfirmed (0conf) funding transaction.
1. LDK rebroadcasts cause problems of their own.
The second bullet mentions relying on a later LDK rebroadcast for repair. Rebroadcasts also cause some of the problems this issue tracks. With the pinned LDK rev (
9174965), once a 0conf splice is locked —splice_lockedgoes out at zero confirmations because a splice inherits the channel'sminimum_depth— LDK rebroadcasts the still-unconfirmed funding transaction on every monitor-update completion until it confirms, including after restarts. These rebroadcasts come typedTransactionType::Funding { channels }rather thanInteractiveFunding, so they go through the generic classification path, with amounts taken from the on-chain wallet rather than from the splice contributions. On current main:
classify_fundingrecords every broadcast unconditionally, without checking for wallet activity. Both peers broadcast the fully-signed transaction, so the side that didn't contribute — whose wallet sees nothing — gets a zero-amount payment record for each rebroadcast.- If the rebroadcast txid matches the record's id (the first candidate), the reclassification has no guard against
Fundinglanding on anInteractiveFundingrecord: the wallet-derived amounts and the less specific type overwrite the negotiated ones.- If it doesn't (a non-first RBF candidate),
classify_fundingfalls back toPaymentId(txid)without consultingfind_payment_by_txidand creates a duplicate. So classification has the same txid-resolution problem the checklist raises for wallet sync.I hit all three on #930's branch; a regression test there (0conf splice, then payments to drive monitor updates) reproduces them. I'm splitting the fixes for the first two, plus the test, out of #930 into a separate PR. The third checklist item should probably also ask what a redelivered broadcast looks like, since it may come with a different type and amounts. It may also be worth asking upstream whether these rebroadcasts could keep the
InteractiveFundingtype — thoughTransactionTypeis documented as best-effort, so we'd still need to handle it.
Opened #1049 for this.
One concrete case from #1024: WalletEvent::TxReplaced reads the authoritative PaymentDetails before refreshing its pending-store mirror. With a bounded payment cache, that read can now reach the KVStore and fail after BDK has already applied the wallet update. update_payment_store then consumes the event and skips the remainder of the batch, while those events may not be generated again.
This is the same recovery gap, but triggered by a read failure rather than a partial write. We should define how wallet-sync events are retained or retried after payment-store failures instead of falling back to potentially stale pending metadata. See the #1024 review thread (#1024 (comment)).
PR #962 addresses the in-process races between wallet sync, funding classification, and payment graduation. Several persistence and identity questions remain deliberately unresolved.
Known remaining problems:
Relevant discussions:
Follow-up questions/checklist:
DataStoreentries in-memory, add pagination #1024’s payment-store caching changes.This issue intentionally does not prescribe a storage or recovery design.