Our startup helper retries more than it knows (inspection-cycle execution)

DavidBennett0780 · 8 Aug 2026, 11:21 UTC

Reply to discussion
DA
DavidBennett0780
Our application crashed after sending a inspection-cycle execution request but before saving the result. Universal Robots UR5e may have continued; our local file still says pending. We're in a workshop fixture inspection cell, using a reference plate. The startup helper resubmits pending jobs automatically. i'm fixing that label because it mixes never submitted with submitted but no known outcome.

15 replies

AN
AnilBrooks0810
Replying to DavidBennett0780

What is the last durable entry for this attempt? Compare it with matching acceptance evidence, without letting startup resend it.

19 points
DA
DavidBennett0780
Replying to AnilBrooks0810

@AnilBrooks0810 Our last durable entry is submission intent. Controller history has matching acceptance, but no surviving completion record. i've kept it out of the resend queue.

6 points
AN
AnilBrooks0810
Replying to DavidBennett0780

@DavidBennett0780 Preserve the uncertain attempt and use matching evidence for reconciliation. An offline crash replay plus a local transaction for state and count should provide atomic recovery.

15 points
DA
DavidAdams0171
Replying to AnilBrooks0810

Atomic locally. That transaction can't include the controller accepting a request. The gap that caused this still exists.

4 points
AN
AnilBrooks0810
Replying to DavidAdams0171

@DavidAdams0171 You're right to distinguish those. I meant atomic local accounting updates, not an atomic controller exchange; the acceptance gap still requires explicit uncertainty and reconciliation.

18 points
DA
DavidBennett0780
Replying to AnilBrooks0810

@AnilBrooks0810 So persist intent first, but don't treat intent as proof of sending? That's where our pending label got stretched beyond usefulness.

5 points
AN
AnilBrooks0810
Replying to DavidBennett0780

Correct: persisting intent records the local decision to submit. Evidence of sending, remote acceptance and completion are separate observations and shouldn't be inferred from it.

19 points
KA
KaiBennett0715
Replying to AnilBrooks0810

Does missing local acceptance mean startup should put the job back in the unsubmitted queue, or can acceptance have occurred without reaching that record?

16 points
AN
AnilBrooks0810
Replying to KaiBennett0715

@KaiBennett0715 Acceptance can happen before your local save. A crash in that gap leaves uncertainty, which is exactly why missing local acceptance can't mean unsubmitted.

10 points
RA
RaviChen1172
Replying to AnilBrooks0810

@AnilBrooks0810 On my ledger, we simulated crashes before send but forgot the gap after remote acceptance. All the restart tests passed for the easy half.

25 points
DA
DavidBennett0780
Replying to RaviChen1172

@RaviChen1172 i'll put our replay break exactly in that gap. Startup should retain uncertainty and emit no replacement request.

13 points
DA
DavidAdams0171
Replying to DavidBennett0780

And what evidence ends uncertainty? A matching acceptance still doesn't tell you the inspection finished.

18 points
DA
DavidBennett0780
Replying to DavidAdams0171

@DavidAdams0171 For completion, matching completion evidence. Our acceptance record only rules out treating this as known unsubmitted work.

6 points
DA
DavidBennett0780
Replying to DavidBennett0780

i've kept this accepted-but-uncertain attempt out of our resend queue. The final outcome and the broader startup fix are still open.

7 points
AN
AnilBrooks0810
Replying to DavidBennett0780

That's a concrete partial result while leaving the final outcome and wider implementation open.

15 points

Add to the discussion

Welcome to Application Robot

Everyone can read the forum. Sign in or create an account to start a discussion, reply, or upload photos.

Forgot your password?

By creating an account, you agree to our Terms and Conditions and community guidelines. Read our Privacy Policy for how your information is handled.