WIP: Introduce an etcd operator leader status field #694

ironcladlou · 2020-07-15T20:05:51Z

This patch explores adding an etcd operator status field which reports
leader member information. Exposing this would allow, for example, smarter
decision-making by the MCO regarding reboot ordering by providing a hint
to help minimize disruption via excessive leader changes.

openshift-ci-robot · 2020-07-15T20:06:07Z

[APPROVALNOTIFIER] This PR is NOT APPROVED

This pull-request has been approved by: ironcladlou
To complete the pull request process, please assign knobunc
You can assign the PR to them by writing /assign @knobunc in a comment when ready.

The full list of commands accepted by this bot can be found here.

Needs approval from an approver in each of these files:

OWNERS

Approvers can indicate their approval by writing /approve in a comment
Approvers can cancel approval by writing /approve cancel in a comment

ironcladlou · 2020-07-15T20:06:54Z

See the discussion in openshift/machine-config-operator#1897 for more details about one way this could be useful.

This patch explores adding an etcd operator status field which reports leader member information. Exposing this would allow, for example, smarter decision-making by the MCO regarding reboot ordering by providing a hint to help minimize disruption via excessive leader changes.

cgwalters · 2020-07-15T20:21:18Z

operator/v1/types_etcd.go

+	// name is the etcd leader member name, if available.
+	Name string `json:"name,omitempty"`
+	// node is the etcd leader member node, if available.
+	Node string `json:"node,omitempty"`


Shouldn't this be e.g. NodeRef *corev1.ObjectReference ?

Shouldn't this be e.g. NodeRef *corev1.ObjectReference ?

No. We encourage the creation of specific reference types. See the reasoning in https://github.com/kubernetes/api/blob/master/core/v1/types.go#L5172-L5186

@cgwalters

Ah, I see. Hmm. So, since there's nothing more we care about here to a node than its name, does that argue for keeping it as a string?

openshift-ci-robot · 2020-07-15T20:21:37Z

@ironcladlou: The following test failed, say /retest to rerun all failed tests:

Test name	Commit	Details	Rerun command
ci/prow/verify	`06b1a7d`	link	`/test verify`

Full PR test history. Your PR dashboard.

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository. I understand the commands that are listed here.

ironcladlou · 2020-07-16T12:12:50Z

P.S., this is for now to help promote discussion and experimentation — we haven't actually done the work/measurement to prove the cited use case has a benefit with expanding the API for

deads2k · 2020-07-16T14:33:54Z

@ironcladlou what about something like

status struct{
  conditions
  []EtcdMembers
}

EtcdMembers struct{
  name string
  node string
  status string // learning,unhealthy,not-a-member,healthy
  memberType string // leader,follower (whatever this is)
  fsyncP99InLastMinute string
  peerLatencyP99InLastMinute string

  // or whatever it is you want
}

cgwalters · 2020-07-20T22:26:50Z

OK so I'm about at the point in openshift/machine-config-operator#1897 work where I'm blocking on this, because I really want to test OS update things after we're only updating etcd followers first.

I starting looking at hacking something into the MCD for this but it's ugly.

cgwalters · 2020-07-23T13:29:32Z

If we don't land this my tentative thoughts are:

Patch the MCD to detect if it's on control plane, if so start watching the local etcd and find out if the current node is a leader
If it is a leader add mco.openshift.io/etcdleader: "" label to node, if not remove label if exists

In playing with this what I'm stumbling over is getting the right creds set up to talk to etcd, and I don't know if there's a non-polling way to watch for leader changes.

The rest of the logic for managing upgrade order would be in the MCC and proceed basically the same.

cgwalters · 2020-07-24T14:33:14Z

You know what may be far simpler (and avoid the "etcd writes its state to kube which writes to etcd" issue) is having the etcd pod write out something extremely simple like /run/etcd/status.json or whatever, then the MCD pod on the control plane can use inotify to watch it. Or we can set up some sort of local read-only communication channel.

Anyways for now I wrote some hacky code that reuses the certs from the apiserver in the MCD
openshift/machine-config-operator#1946

cgwalters · 2020-07-30T20:22:18Z

Or, rather than trying to extend the API here we could use a node annotation.

openshift-bot · 2020-11-01T01:36:57Z

Issues go stale after 90d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle stale.
Stale issues rot after an additional 30d of inactivity and eventually close.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle stale

openshift-bot · 2020-12-01T03:33:07Z

Stale issues rot after 30d of inactivity.

Mark the issue as fresh by commenting /remove-lifecycle rotten.
Rotten issues close after an additional 30d of inactivity.
Exclude this issue from closing by commenting /lifecycle frozen.

If this issue is safe to close now please do so with /close.

/lifecycle rotten
/remove-lifecycle stale

openshift-bot · 2020-12-31T05:22:39Z

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

openshift-ci-robot · 2020-12-31T05:22:53Z

@openshift-bot: Closed this PR.

In response to this:

Rotten issues close after 30d of inactivity.

Reopen the issue by commenting /reopen.
Mark the issue as fresh by commenting /remove-lifecycle rotten.
Exclude this issue from closing again by commenting /lifecycle frozen.

/close

Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes/test-infra repository.

openshift-ci-robot added the do-not-merge/work-in-progress Indicates that a PR should not merge because it is a work in progress. label Jul 15, 2020

openshift-ci-robot requested review from jwforres and smarterclayton July 15, 2020 20:06

ironcladlou mentioned this pull request Jul 15, 2020

Bug 1850057: stage OS updates (nicely) while etcd is still running openshift/machine-config-operator#1897

Closed

ironcladlou force-pushed the etcd-leader-status branch from 6ffe8b6 to 06b1a7d Compare July 15, 2020 20:14

cgwalters reviewed Jul 15, 2020

View reviewed changes

cgwalters mentioned this pull request Jul 30, 2020

[WIP] Bug 1850057: update etcd followers first, use bfq on control plane openshift/machine-config-operator#1946

Closed

openshift-ci-robot added the lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. label Nov 1, 2020

openshift-ci-robot added lifecycle/rotten Denotes an issue or PR that has aged beyond stale and will be auto-closed. and removed lifecycle/stale Denotes an issue or PR has remained open with no activity and has become stale. labels Dec 1, 2020

openshift-ci-robot closed this Dec 31, 2020

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

WIP: Introduce an etcd operator leader status field #694

WIP: Introduce an etcd operator leader status field #694

ironcladlou commented Jul 15, 2020

openshift-ci-robot commented Jul 15, 2020

ironcladlou commented Jul 15, 2020

cgwalters Jul 15, 2020 •

edited

Loading

deads2k Jul 16, 2020

cgwalters Jul 17, 2020

openshift-ci-robot commented Jul 15, 2020

ironcladlou commented Jul 16, 2020

deads2k commented Jul 16, 2020

cgwalters commented Jul 20, 2020

cgwalters commented Jul 23, 2020

cgwalters commented Jul 24, 2020

cgwalters commented Jul 30, 2020

openshift-bot commented Nov 1, 2020

openshift-bot commented Dec 1, 2020

openshift-bot commented Dec 31, 2020

openshift-ci-robot commented Dec 31, 2020

WIP: Introduce an etcd operator leader status field #694

WIP: Introduce an etcd operator leader status field #694

Conversation

ironcladlou commented Jul 15, 2020

openshift-ci-robot commented Jul 15, 2020

ironcladlou commented Jul 15, 2020

cgwalters Jul 15, 2020 • edited Loading

Choose a reason for hiding this comment

deads2k Jul 16, 2020

Choose a reason for hiding this comment

cgwalters Jul 17, 2020

Choose a reason for hiding this comment

openshift-ci-robot commented Jul 15, 2020

ironcladlou commented Jul 16, 2020

deads2k commented Jul 16, 2020

cgwalters commented Jul 20, 2020

cgwalters commented Jul 23, 2020

cgwalters commented Jul 24, 2020

cgwalters commented Jul 30, 2020

openshift-bot commented Nov 1, 2020

openshift-bot commented Dec 1, 2020

openshift-bot commented Dec 31, 2020

openshift-ci-robot commented Dec 31, 2020

cgwalters Jul 15, 2020 •

edited

Loading