[JENKINS-36013] Automatically abort ExecutorPickle rehydration from an ephemeral node - Jenkins Jira

Type: Improvement
Resolution: Fixed
Priority: Critical
Component/s: workflow-durable-task-step-plugin
Labels:
- cloudbees-internal-pipeline
- robustness

Similar Issues:
Powered by SuggestiMate

Show
Epic Link:
Pipeline Durability
Sprint:
Pipeline - July/August

ExecutorPickle.rehydrate ought to be able to detect that it has been spinning in circles because the agent node it was supposed to run on is not in the Jenkins node list, and automatically abort, causing the build to fail with a comprehensible message rather than just hanging indefinitely. (As opposed to being registered but offline, which is normal enough for a JNLP agent etc.—in such cases we just want to wait for the agent to come back online.)

This would provide a better experience for the case of a build which was running on an EphemeralNode (such as from a Cloud without durable-task integration) when Jenkins was restarted. An agent using an inappropriate RetentionStrategy is trickier since it might still be defined after a restart, but will soon be terminated. Similarly, there may be cases where the agent is actually going to be redefined (with the same name) when it is attached after the restart—not sure about the Swarm plugin, but CloudBees DEV@cloud OPEs work this way. To prevent the build from being killed too aggressively, the cleanup should be delayed until some time has elapsed since rehydration began (or, ideally, since Jenkins completed initialization)—say, five minutes.

depends on

JENKINS-26130 Print progress of pending pickles

Resolved

is duplicated by

JENKINS-45917 [Jenkins v2.63] Build queue deadlocks

Closed

is related to

JENKINS-45917 [Jenkins v2.63] Build queue deadlocks

Closed

JENKINS-41569 Pipeline hangs waiting for resume on an agent which never was

Closed

relates to

JENKINS-41569 Pipeline hangs waiting for resume on an agent which never was

Closed

JENKINS-49707 Auto retry for elastic agents after channel closure

Resolved

JENKINS-43607 Jenkins pipeline not aborted when the machine running docker container goes offline

Resolved

JENKINS-33761 Ability to disable Pipeline durability and "resume" build.

Closed

links to

CloudBees Internal CD-179

CloudBees Internal CLTS-2226

workflow-durable-task-step #47

workflow-durable-task-step #48

(3 relates to, 4 links to)

Assignee:: Sam Van Oort

Reporter:: Jesse Glick

Votes:: 6 Vote for this issue

Watchers:: 11 Start watching this issue

Created:: 2016-06-16 14:12

Updated:: 2019-04-29 18:52

Resolved:: 2017-08-23 22:22

Details

Description

Attachments

Issue Links

Activity

People

Dates