typhoon

Commit Graph

Author	SHA1	Message	Date
Dalton Hubble	8f0d2b5db4	Update Grafana from v5.2.4 to v5.3.0	2018-10-13 23:03:31 -07:00
Dalton Hubble	2e89e161e9	Remove Azure admin_password (disabled) now that its optional * Requires terraform-provider-azurerm v1.16.0 or higher https://github.com/terraform-providers/terraform-provider-azurerm/pull/1958	2018-10-13 22:40:58 -07:00
Dalton Hubble	55bb4dfba6	Raise CoreDNS replica count to 2 or more * Run at least two replicas of CoreDNS to better support rolling updates (previously, kube-dns had a pod nanny) * On multi-master clusters, set the CoreDNS replica count to match the number of masters (e.g. a 3-master cluster previously used replicas:1, now replicas:3) * Add AntiAffinity preferred rule to favor distributing CoreDNS pods across controller nodes nodes	2018-10-13 20:31:29 -07:00
Dalton Hubble	43fe78a2cc	Raise scheduler/controller-manager replicas in multi-master * Continue to ensure scheduler and controller-manager run at least two replicas to support performing kubectl edits on single-master clusters (no change) * For multi-master clusters, set scheduler / controller-manager replica count to the number of masters (e.g. a 3-master cluster previously used replicas:2, now replicas:3)	2018-10-13 16:16:29 -07:00
Dalton Hubble	5a283b6443	Update etcd from v3.3.9 to v3.3.10 * https://github.com/etcd-io/etcd/blob/master/CHANGELOG-3.3.md#v3310-2018-10-10	2018-10-13 13:14:37 -07:00
Dalton Hubble	db36036c81	Require terraform-provider-digitalocean plugin ~> 1.0 * Require a terraform-provider-digitalocean plugin version of 1.0 or higher within the same major version (e.g. allow 1.1 but not 2.0) * Change requirement from ~> 0.1.2 (which allowed up to but not including 1.0 release)	2018-10-02 17:09:19 +02:00
Dalton Hubble	7653e511be	Update CoreDNS and Calico versions * Update CoreDNS from 1.1.3 to 1.2.2 * Update Calico from v3.2.1 to v3.2.3	2018-10-02 16:07:48 +02:00
Dalton Hubble	032a24133b	Update Prometheus from v2.3.2 to v2.4.2 * https://github.com/prometheus/prometheus/releases/tag/v2.4.0 * https://github.com/prometheus/prometheus/releases/tag/v2.4.1 * https://github.com/prometheus/prometheus/releases/tag/v2.4.2	2018-09-21 22:27:11 -07:00
Dalton Hubble	ad871dbfa9	Update Kubernetes from v1.11.2 to v1.11.3 * https://github.com/kubernetes/kubernetes/blob/master/CHANGELOG-1.11.md#v1113	2018-09-13 18:50:41 -07:00
Dalton Hubble	dc03f7a4a9	Update nginx-ingress from 0.17.1 to 0.19.0 * If using --enable-ssl-passthrough or exposing TCP/UDP services, be aware of https://github.com/kubernetes/ingress-nginx/pull/3038 * Workarounds until the fix merges are to stay on 0.17.1, use the suggested development image, or revert to securityContext `runAsNonRoot: false` for a while (less secure)	2018-09-08 17:57:01 -07:00
Dalton Hubble	1b8234eb91	Update Grafana from v5.2.2 to v5.2.4 * https://github.com/grafana/grafana/releases/tag/v5.2.3 * https://github.com/grafana/grafana/releases/tag/v5.2.4	2018-09-08 15:41:20 -07:00
Dalton Hubble	4ba090feb0	Update kube-state-metrics from v1.3.1 to v1.4.0	2018-08-29 09:37:50 -07:00
Dalton Hubble	4882fe1053	Add docs for Azure Ingress and worker pools * Azure worker pools must be in the same region as the cluster itself unfortunately	2018-08-27 23:30:56 -07:00
Dalton Hubble	7eb09237f4	Update Calico from v3.1.3 to v3.2.1 * Add new bird and felix readiness checks * Read MTU from ConfigMap veth_mtu * Add RBAC read for serviceaccounts * Remove invalid description from CRDs	2018-08-25 17:53:11 -07:00
Dalton Hubble	e58b424882	Fix firewall to allow etcd client traffic between controllers * Broaden internal-etcd firewall rule to allow etcd client traffic (2379) from other controller nodes * Previously, kube-apiservers were only able to connect to their node's local etcd peer. While master node outages were tolerated, reaching a healthy peer took longer than neccessary in some cases * Reduce time needed to bootstrap a cluster	2018-08-21 23:51:40 -07:00
Dalton Hubble	ea365b551a	Fix docs mentions of ELBs to NLBs * Typhoon AWS clusters use an NLB rather than an ELB, since v1.10.5 * Add a few missing links in CHANGES	2018-08-21 21:40:06 -07:00
Dalton Hubble	bbf2c13eef	Remove AWS security rule allowing ICMP packets to nodes * Deny ICMP packets for consistency across Typhoon clusters on various clouds and because there isn't much need to allow them	2018-08-21 21:16:16 -07:00
Dalton Hubble	da5d2c5321	Remove GCP firewall rule allowing Nginx Ingress health * Nginx Ingress addon no longer uses hostNework so Prometheus may scrape port 10254 via the CNI network, rather than via the host address	2018-08-21 21:06:03 -07:00
Dalton Hubble	bec5250e73	Remove unofficial bare-metal _networkds variables Remove controller_networkds and worker_networkds variables. These variables were always listed as experimental, unsupported, and excluded from documentation in anticipation of Container Linux Config snippets * Use Container Linux Config snippets on bare-metal instead. They provide safer, more powerful, and more elegant host customization	2018-08-13 23:33:29 -07:00
Dalton Hubble	dbdc3fc850	Add nginx-ingress addon manifests for bare-metal	2018-08-11 12:14:23 -07:00
Dalton Hubble	e00f97c578	Update nginx-ingress from 0.16.2 to 0.17.1 * https://github.com/kubernetes/ingress-nginx/releases/tag/nginx-0.17.1 * https://github.com/kubernetes/ingress-nginx/releases/tag/nginx-0.17.0	2018-08-08 00:45:20 -07:00
Dalton Hubble	f7ebdf475d	Update Kubernetes from v1.11.1 to v1.11.2 * https://github.com/kubernetes/kubernetes/blob/master/CHANGELOG-1.11.md#v1112	2018-08-07 21:57:25 -07:00
Dalton Hubble	edc250d62a	Fix Kublet version for Fedora Atomic modules * Release v1.11.1 erroneously left Fedora Atomic clusters using the v1.11.0 Kubelet. The rest of the control plane ran v1.11.1 as expected * Update Kubelet from v1.11.0 to v1.11.1 so Fedora Atomic matches Container Linux * Container Linux modules were not affected	2018-07-29 12:13:29 -07:00
Dalton Hubble	db64ce3312	Update etcd from v3.3.8 to v3.3.9 * https://github.com/coreos/etcd/blob/master/CHANGELOG-3.3.md#v339-2018-07-24	2018-07-29 11:27:37 -07:00
Dalton Hubble	7c327b8bf4	Update from bootkube v0.12.0 to v0.13.0	2018-07-29 11:20:17 -07:00
Dalton Hubble	e6720cf738	Update heapster from v1.5.3 to v1.5.4 * https://github.com/kubernetes/heapster/releases/tag/v1.5.4	2018-07-29 11:19:57 -07:00
Dalton Hubble	844f380b4e	Update Grafana from v5.2.1 to v5.2.2 * https://github.com/grafana/grafana/releases/tag/v5.2.2	2018-07-29 11:12:56 -07:00
Dalton Hubble	4e7dfc115d	Support Container Linux Config snippets on bare-metal	2018-07-25 23:14:54 -07:00
Dalton Hubble	d8d524d10b	Update Kubernetes from v1.11.0 to v1.11.1 * https://github.com/kubernetes/kubernetes/blob/master/CHANGELOG-1.11.md#v1111	2018-07-20 00:41:27 -07:00
Dalton Hubble	02cd8eb8d3	Update Prometheus from v2.3.1 to v2.3.2 * https://github.com/prometheus/prometheus/releases/tag/v2.3.2	2018-07-14 14:25:49 -07:00
Dalton Hubble	3352388fe6	Update changelog and docs for release	2018-07-04 12:28:25 -07:00
Dalton Hubble	915f89d3c8	Update Fedora Atomic from 27 to 28 on bare-metal	2018-07-04 11:41:54 -07:00
Dalton Hubble	f40f60b83c	Update Nginx Ingress controller from 0.15.0 to 0.16.2 * https://github.com/kubernetes/ingress-nginx/releases/tag/nginx-0.16.2 * https://github.com/kubernetes/ingress-nginx/blob/master/Changelog.md	2018-07-02 22:06:22 -07:00
Dalton Hubble	6f958d7577	Replace kube-dns with CoreDNS * Add system:coredns ClusterRole and binding * Annotate CoreDNS for Prometheus metrics scraping * Remove kube-dns deployment, service, & service account * https://github.com/poseidon/terraform-render-bootkube/pull/71 * https://kubernetes.io/blog/2018/06/27/kubernetes-1.11-release-announcement/	2018-07-01 22:55:01 -07:00
Dalton Hubble	ee31074679	Promote Typhoon Google Cloud for Container Linux to stable	2018-07-01 22:52:27 -07:00
Dalton Hubble	18502d64d6	Update Fedora Atomic from 27 to 28 on GCP	2018-07-01 22:46:51 -07:00
Dalton Hubble	a3349b5c68	Update heapster from v1.5.2 to v1.5.3	2018-07-01 21:07:52 -07:00
Dalton Hubble	74dc6b0bf9	Update Grafana from 5.1.4 to 5.2.1 * http://docs.grafana.org/guides/whats-new-in-v5-2/ * https://github.com/grafana/grafana/releases/tag/v5.2.0 * https://github.com/grafana/grafana/releases/tag/v5.2.1	2018-07-01 20:55:34 -07:00
Dalton Hubble	fd1de27aef	Remove deprecated ingress_static_ip and controllers_ipv4_public outputs	2018-07-01 20:47:46 -07:00
Dalton Hubble	93de7506ef	Update Fedora Atomic from 27 to 28 on AWS	2018-06-30 18:55:18 -07:00
Dalton Hubble	8464b258d8	Update Kubernetes from v1.10.5 to v1.11.0 * Force apiserver to stop listening on 127.0.0.1:8080 * Remove deprecated Kubelet `--allow-privileged`. Defaults to true. Use `PodSecurityPolicy` if limiting is desired * https://github.com/kubernetes/kubernetes/releases/tag/v1.11.0 * https://github.com/poseidon/terraform-render-bootkube/pull/68	2018-06-27 22:47:35 -07:00
Dalton Hubble	855aec5af3	Clarify AWS module output names and changes	2018-06-23 15:29:13 -07:00
Dalton Hubble	0c4d59db87	Use global HTTP/TCP proxy load balancing for Ingress on GCP * Switch Ingress from regional network load balancers to global HTTP/TCP Proxy load balancing * Reduce cost by ~$19/month per cluster. Google bills the first 5 global and regional forwarding rules separately. Typhoon clusters now use 3 global and 0 regional forwarding rules. * Worker pools no longer include an extraneous load balancer. Remove worker module's `ingress_static_ip` output. * Add `ingress_static_ipv4` output variable * Add `worker_instance_group` output to allow custom global load balancing * Deprecate `controllers_ipv4_public` module output * Deprecate `ingress_static_ip` module output. Use `ingress_static_ipv4`	2018-06-23 14:37:40 -07:00
Dalton Hubble	2eaf04c68b	Drop hostNetwork from nginx-ingress addon * Both flannel and Calico support host port via `portmap` * Allows writing NetworkPolicies that reference ingress pods in `from` or `to`. HostNetwork pods were difficult to write network policy for since they could circumvent the CNI network to communicate with pods on the same node.	2018-06-22 00:46:41 -07:00
Dalton Hubble	fb6f40051f	Disable AWS detailed monitoring on worker nodes * Basic monitoring (free) is sufficient for casual console browsing * Detailed monitoring (paid) is not leveraged for CloudWatch anyway * Favor Prometheus for cloud-agnostic metrics, aggregation, and alerting	2018-06-22 00:26:06 -07:00
Dalton Hubble	316f06df06	Combine NLBs to use one NLB per cluster * Simplify clusters to come with a single NLB * Listen for apiserver traffic on port 6443 and forward to controllers (with healthy apiserver) * Listen for ingress traffic on ports 80/443 and forward to workers (with healthy ingress controller) * Reduce cost of default clusters by 1 NLB ($18.14/month) * Keep using CNAME records to the `ingress_dns_name` NLB and the nginx-ingress addon for Ingress (up to a few million RPS) * Users with heavy traffic (many million RPS) can create their own separate NLB(s) for Ingress and use the new output worker target groups * Fix issue where additional worker pools come with an extraneous network load balancer	2018-06-21 23:46:57 -07:00
Dalton Hubble	f4d3059b00	Update Kubernetes from v1.10.4 to v1.10.5 * https://github.com/kubernetes/kubernetes/blob/master/CHANGELOG-1.10.md#v1105	2018-06-21 22:51:39 -07:00
Dalton Hubble	6c5a1964aa	Change kube-apiserver port from 443 to 6443 * Adjust firewall rules, security groups, cloud load balancers, and generated kubeconfig's * Facilitates some future simplifications and cost reductions * Bare-Metal users who exposed kube-apiserver on a WAN via their router or load balancer will need to adjust its configuration. This is uncommon, most apiserver are on LAN and/or behind VPN so no routing infrastructure is configured with the port number	2018-06-19 23:48:51 -07:00
Dalton Hubble	6e64634748	Update etcd from v3.3.7 to v3.3.8 * https://github.com/coreos/etcd/releases/tag/v3.3.8	2018-06-19 21:56:21 -07:00
Dalton Hubble	d5de41e07a	Update Grafana from 5.1.3 to 5.1.4 * https://github.com/grafana/grafana/releases/tag/v5.1.4	2018-06-19 21:45:15 -07:00
Dalton Hubble	05b99178ae	Update prometheus from v2.3.0 to v2.3.1 * https://github.com/prometheus/prometheus/releases/tag/v2.3.1	2018-06-19 21:43:50 -07:00
Dalton Hubble	ed0b781296	Fix possible deadlock for provisioning bare-metal clusters * Closes #235	2018-06-14 23:15:28 -07:00
Dalton Hubble	51906bf398	Update etcd from v3.3.6 to v3.3.7	2018-06-14 22:46:16 -07:00
Stephen Demos	18dd7ccc09	Update CLUO from v0.6.0 to v0.7.0	2018-06-14 22:32:36 -07:00
Dalton Hubble	cbe646fba6	Label namespaces to ease writing Network Policies	2018-06-09 11:45:11 -07:00
Dalton Hubble	c166b2ba33	Update prometheus from v2.2.1 to v2.3.0	2018-06-09 11:43:10 -07:00
Dalton Hubble	79260c48f6	Update Kubernetes from v1.10.3 to v1.10.4	2018-06-06 23:23:11 -07:00
Dalton Hubble	589c3569b7	Update etcd from v3.3.5 to v3.3.6 * https://github.com/coreos/etcd/releases/tag/v3.3.6	2018-06-06 23:19:30 -07:00
Dalton Hubble	d32e6797ae	Annotate Grafana so Prometheus scrapes metrics	2018-05-30 22:37:47 -07:00
Dalton Hubble	32a9a83190	Add Prometheus liveness and readiness probes	2018-05-30 22:34:07 -07:00
Dalton Hubble	6e968cd152	Update Calico from v3.1.2 to v3.1.3 * https://github.com/projectcalico/calico/releases/tag/v3.1.3 * https://github.com/projectcalico/cni-plugin/releases/tag/v3.1.3	2018-05-30 21:32:12 -07:00
Dalton Hubble	4ea1fde9c5	Update Kubernetes from v1.10.2 to v1.10.3 * https://github.com/kubernetes/kubernetes/blob/master/CHANGELOG-1.10.md#v1103 * Update Calico from v3.1.1 to v3.1.2	2018-05-21 21:38:43 -07:00
Dalton Hubble	1e2eec6487	Update Fedora Atomic from 27 to 28 on DigitalOcean * Fedora Atomic 27 images disappeared from DigitalOcean and forced this early update (there are known bugs)	2018-05-21 21:30:23 -07:00
Dalton Hubble	28d0891729	Annotate nginx-ingress addon for Prometheus auto-discovery * Add Google Cloud firewall rule to allow worker to worker access to health and metrics	2018-05-19 13:13:14 -07:00
Dalton Hubble	714419342e	Update nginx-ingress from 0.14.0 to 0.15.0 * https://github.com/kubernetes/ingress-nginx/releases/tag/nginx-0.15.0	2018-05-17 21:42:55 -07:00
Dalton Hubble	3701c0b1fe	Update Grafana from v5.1.2 to v5.1.3 * https://github.com/grafana/grafana/releases/tag/v5.1.3	2018-05-17 21:36:09 -07:00
Dalton Hubble	0c3557e68e	Allow Flatcar Linux os_channel on bare-metal * Choose the Container Linux derivative Flatcar Linux on bare-metal by setting os_channel to flatcar-stable, flatcar-beta or flatcar-alpha * As with Container Linux from Red Hat, the version (os_version) must correspond to the channel being used * Thank you to @dongsupark from Kinvolk	2018-05-17 20:09:36 -07:00
Dalton Hubble	adc6c6866d	Rename container_linux_ bare-metal variables * Allow for Container Linux derivatives * Replace container_linux_channel variable with `os_channel` * Replace `container_linux_version` variable with `os_version` * Please change values `stable`, `beta`, or `alpha` to `coreos-stable`, `coreos-beta`, `coreos-alpha` (action required!)	2018-05-16 22:40:39 -07:00
Dalton Hubble	9ac7b0655f	Add bare-metal network_ip_autodetection_method variable for multi-NIC * Allow setting the Calico host IPv4 address autodetection method * Use Calico's default "first-found" method to support single NIC and bonded NIC nodes * Allow methods like `can-reach=IP` or `interface=REGEX` for multi NIC nodes * https://docs.projectcalico.org/v3.1/reference/node/configuration#ip-autodetection-methods	2018-05-15 23:27:34 -07:00
Dalton Hubble	c2b719dc75	Configure Prometheus to scrape Kubelets directly * Use Kubelet bearer token authn/authz to scrape metrics * Drop RBAC permission from nodes/proxy to nodes/metrics * Stop proxying kubelet scrapes through the apiserver, since this required higher privilege (nodes/proxy) and can add load to the apiserver on large clusters	2018-05-14 23:06:50 -07:00
Dalton Hubble	37981f9fb1	Allow bearer token authn/authz to the Kubelet * Require Webhook authorization to the Kubelet * Switch apiserver X509 client cert org to systems:masters to grant the apiserver admin and satisfy the authorization requirement. kubectl commands like logs or exec that have the apiserver make requests of a kubelet continue to work as before * https://kubernetes.io/docs/admin/kubelet-authentication-authorization/ * https://github.com/poseidon/typhoon/issues/215	2018-05-13 23:20:42 -07:00
Dalton Hubble	5eb11f5104	Allow Flatcar Linux os_image on AWS, rename os_channel * Replace os_channel variable with os_image to align naming across clouds. Users who set this option to stable, beta, or alpha should now set os_image to coreos-stable, coreos-beta, or coreos-alpha. * Default os_image to coreos-stable. This continues to use the most recent image from the stable channel as always. * Allow Container Linux derivative Flatcar Linux by setting os_image to `flatcar-stable`, `flatcar-beta`, `flatcar-alpha`	2018-05-12 11:41:58 -07:00
Dalton Hubble	f2ee75ac98	Require Terraform v0.11.x, drop v0.10.x support * Raise minimum Terraform version to v0.11.0 * Terraform v0.11.x has been supported since Typhoon v1.9.2 and Terraform v0.10.x was last released in Nov 2017. I'd like to stop worrying about v0.10.x and remove migration docs as a later followup * Migration docs docs/topics/maintenance.md#terraform-v011x	2018-05-10 02:20:46 -07:00
Dalton Hubble	8b8e364915	Update etcd from v3.3.4 to v3.3.5 * https://github.com/coreos/etcd/releases/tag/v3.3.5	2018-05-10 02:12:53 -07:00
Dalton Hubble	fb88113523	Disable default Google Analytics in Grafana addon * Its come to my attention Grafana reports analytics data by default. Typhoon's philosophy requires user permission for data collection so the addon should have this disabled * http://docs.grafana.org/installation/configuration/#analytics	2018-05-10 01:18:47 -07:00
Dalton Hubble	1854f5c104	Update Grafana from v5.1.1 to v5.1.2 * https://github.com/grafana/grafana/releases/tag/v5.1.2	2018-05-10 01:09:08 -07:00
Dalton Hubble	726b58b697	Update Grafana from v5.0.4 to v5.1.1 * https://github.com/grafana/grafana/releases/tag/v5.1.1 * https://github.com/grafana/grafana/releases/tag/v5.1.0	2018-05-07 22:05:19 -07:00
Dalton Hubble	a54e3c0da1	Fix Prometheus data dir to /var/lib/prometheus * A data volume (emptyDir) is mounted to /var/lib/prometheus * Users could swap emptyDir for any desired volume if data persistence is desired. Prometheus previously defaulted to keeping its data in ./data relative to /prometheus. Override this behavior to store data in /var/lib/prometheus	2018-05-01 22:05:27 -07:00
Dalton Hubble	cc29530ba0	Allow preemptible workers on AWS via spot instances * Add `worker_price` to allow worker spot instances. Defaults to empty string for the worker autoscaling group to use regular on-demand instances. * Add `spot_price` to internal `workers` module for spot worker pools * Note: Unlike GCP `preemptible` workers, spot instances require you to pick a bid price.	2018-04-29 13:31:17 -07:00
Dalton Hubble	385584b712	Add changelog notes for release	2018-04-29 12:04:44 -07:00
Dalton Hubble	32ddfa94e1	Update Kubernetes from v1.10.1 to v1.10.2 * https://github.com/kubernetes/kubernetes/releases/tag/v1.10.2	2018-04-28 00:27:00 -07:00
Dalton Hubble	681450aa0d	Update etcd from v3.3.3 to v3.3.4 * https://github.com/coreos/etcd/releases/tag/v3.3.4	2018-04-27 23:57:26 -07:00
Dalton Hubble	fafa028052	Add Typhoon for Fedora Atomic to changelog	2018-04-27 23:55:59 -07:00
Dalton Hubble	a54f76db2a	Update Calico from v3.0.4 to v3.1.1 * https://github.com/projectcalico/calico/releases/tag/v3.1.1 * https://github.com/projectcalico/calico/releases/tag/v3.1.0	2018-04-21 18:30:36 -07:00
Dalton Hubble	e0d9e9979c	Update nginx-ingress from 0.12.0 to 0.13.0 * https://github.com/kubernetes/ingress-nginx/releases/tag/nginx-0.13.0	2018-04-18 21:12:09 -07:00
Dalton Hubble	ad2e4311d1	Switch GCP network lb to global TCP proxy lb * Allow multi-controller clusters on Google Cloud * GCP regional network load balancers have a long open bug in which requests originating from a backend instance are routed to the instance itself, regardless of whether the health check passes or not. As a result, only the 0th controller node registers. We've recommended just using single master GCP clusters for a while * https://issuetracker.google.com/issues/67366622 * Workaround issue by switching to a GCP TCP Proxy load balancer. TCP proxy lb routes traffic to a backend service (global) of instance group backends. In our case, spread controllers across 3 zones (all regions have 3+ zones) and organize them in 3 zonal unmanaged instance groups that serve as backends. Allows multi-controller cluster creation * GCP network load balancers only allowed legacy HTTP health checks so kubelet 10255 was checked as an approximation of controller health. Replace with TCP apiserver health checks to detect unhealth or unresponsive apiservers. * Drawbacks: GCP provision time increases, tailed logs now timeout (similar tradeoff in AWS), controllers only span 3 zones instead of the exact number in the region * Workaround in Typhoon has been known and posted for 5 months, but there still appears to be no better alternative. Its probably time to support multi-master and accept the downsides	2018-04-18 00:09:06 -07:00
Dalton Hubble	9789881243	Update kube-state-metrics from v1.3.0 to v1.3.1 * https://github.com/kubernetes/kube-state-metrics/releases/tag/v1.3.1	2018-04-15 17:10:02 -07:00
Dalton Hubble	77c0a4cf2e	Update Kubernetes from v1.10.0 to v1.10.1 * Use kubernetes-incubator/bootkube v0.12.0	2018-04-12 20:57:31 -07:00
Dalton Hubble	5035d56db2	Refactor GCP to remove controller internal module * Remove the controller internal module to align with other platforms and since its not a supported use case	2018-04-12 19:41:51 -07:00
Dalton Hubble	d276fffcda	Fix bare-metal multiple apply/ssh on Terraform v0.11.4+ * Terraform v0.11.4 introduced changes to remote-exec that mean Typhoon bare-metal clusters require multiple runs of terraform apply to ssh and bootstrap. * Bare-metal installs PXE boot a live instance to install to disk and then reboot from disk as controllers/workers. Terraform remote-exec has no way to "know" to wait until the reboot has occurred to kickoff Kubernetes bootstrap. Previously Typhoon created a "debug" user during this install phase to allow an admin to SSH, but remote-exec would hang, trying to connect as user "core". Terraform v0.11.4 changes this behavior so remote-exec fails and a user must re-run terraform apply until succeeding. * A new way to "trick" remote-exec into waiting for the reboot into the disk install is to run SSH on a non-standard port during the disk install. This retains the ability for an admin to SSH during install (most distros don't have this) and fixes the issue so only a single run of terraform apply is needed. * https://github.com/hashicorp/terraform/pull/17359#issuecomment-376415464	2018-04-08 13:32:31 -07:00
Dalton Hubble	6b08bde479	Use k8s.gcr.io instead of gcr.io/google_containers * Kubernetes recommends using the alias to fetch images from the nearest GCR regional mirror, to abstract the use of GCR, and to drop names containing 'google' * https://groups.google.com/forum/#!msg/kubernetes-dev/ytjk_rNrTa0/3EFUHvovCAAJ	2018-04-08 12:57:52 -07:00
Dalton Hubble	7186aa46da	Update kube-state-metrics from v1.2.0 to v1.3.0 * https://github.com/kubernetes/kube-state-metrics/pull/412 * https://github.com/kubernetes/kube-state-metrics/pull/413	2018-04-04 21:04:13 -07:00
Dalton Hubble	18dbaf74ce	Update kube-dns from v1.14.8 to v1.14.9 * https://github.com/kubernetes/kubernetes/pull/61908	2018-04-04 21:00:23 -07:00
Dalton Hubble	ce001e9d56	Update etcd from v3.3.2 to v3.3.3 * https://github.com/coreos/etcd/releases/tag/v3.3.3	2018-04-04 20:32:24 -07:00
Dalton Hubble	d770393dbc	Add etcd metrics, Prometheus scrapes, and Grafana dash * Use etcd v3.3 --listen-metrics-urls to expose only metrics data via http://0.0.0.0:2381 on controllers * Add Prometheus discovery for etcd peers on controller nodes * Temporarily drop two noisy Prometheus alerts	2018-04-03 20:31:00 -07:00
Dalton Hubble	642f7ec22f	Update CHANGES.md with Kubernetes link	2018-03-30 23:12:38 -07:00
Dalton Hubble	f8e9bfb1c0	Add disk_type variable for EBS volume type on AWS * Change EBS volume type from `standard` ("prior generation) to `gp2`. Prometheus alerts are tuned for SSDs * Other platforms have fast enough disks by default	2018-03-29 22:51:54 -07:00
Dalton Hubble	b1e41dcb99	addons: Update from Grafana v4.6.3 to v5.0.4 This reverts commit `c59a9c66b1`.	2018-03-28 19:45:19 -07:00
Dalton Hubble	cfd603bea2	Ensure etcd secrets are only distributed to controller hosts * Previously, etcd secrets were erroneously distributed to worker nodes (permissions 500, ownership etc:etcd).	2018-03-25 23:46:44 -07:00
Dalton Hubble	fdb543e834	Add optional controller_type and worker_type vars on GCP * Remove optional machine_type variable on Google Cloud * Use controller_type and worker_type instead	2018-03-25 22:11:18 -07:00

1 2 3 4 5

235 Commits