postgres-operator

Commit Graph

Author	SHA1	Message	Date
Oleksii Kliukin	2bb7e98268	update individual role secrets from infrastructure roles (#206 ) * Track origin of roles. * Propagate changes on infrastructure roles to corresponding secrets. When the password in the infrastructure role is updated, re-generate the secret for that role. Previously, the password for an infrastructure role was always fetched from the secret, making any updates to such role a no-op after the corresponding secret had been generated.	2018-02-23 17:24:04 +01:00
Oleksii Kliukin	f18bb6eaaa	Make errors in the cluster list function visible. Sometimes the operator does not pick up clusters right away when they are created. The change attempts to shed light on the reason behind that.	2018-02-22 16:45:10 +01:00
Sergey Dudoladov	66a3b6830e	Call fatalf if namespace to watch does not exist	2018-02-20 16:13:48 +01:00
Sergey Dudoladov	dcfc9925f6	Respond to code review	2018-02-20 14:43:02 +01:00
Sergey Dudoladov	e3d2434420	Use '*' as an alias to denote all namespaces	2018-02-16 15:20:26 +01:00
Sergey Dudoladov	088bf70e7d	Merge branch 'master' into support-many-namespaces	2018-02-16 15:06:10 +01:00
Sergey Dudoladov	ec7de38f9b	Make operator watch its own namespace instead of controller's one	2018-02-16 14:22:38 +01:00
Sergey Dudoladov	5e9a21456e	Remove the incorrect service account check	2018-02-15 16:33:53 +01:00
Sergey Dudoladov	155ae8d50f	Rename the function that checks service account existence	2018-02-15 11:14:13 +01:00
Sergey Dudoladov	d5d15b7546	Look for secrets in the deployed namespace	2018-02-14 15:37:30 +01:00
Sergey Dudoladov	06fd9e33f5	Watch the namespace where operator deploys to unless told otherwise	2018-02-13 18:17:47 +01:00
Sergey Dudoladov	4c23917d42	Watch all namespaces if the relevant param is empty string / 'default' if param is unset	2018-02-12 11:47:56 +01:00
Sergey Dudoladov	066f11cbbd	Streamline handling of the watched_namespace param/envvar	2018-02-09 11:39:56 +01:00
Sergey Dudoladov	b5b0b027f2	Handle watched namespace set in operator config map	2018-02-08 14:51:45 +01:00
Sergey Dudoladov	86807d21ba	Kill operator if the namespace to watch does not exist	2018-02-08 14:24:47 +01:00
Sergey Dudoladov	794feee3e1	Fix the bug with the operator always listening to all namespaces	2018-02-08 13:49:44 +01:00
Sergey Dudoladov	de2a028592	Warn if the watched namespace does not exist	2018-02-07 17:43:05 +01:00
Sergey Dudoladov	74fa7b9492	Restrict operator to single watched namespace via env var	2018-02-07 16:44:49 +01:00
Sergey Dudoladov	f194a2ae5a	Introduce changes from the PR #200 by @alexeyklyukin	2018-02-07 14:02:32 +01:00
Sergey Dudoladov	74a1e9661b	Remove setting the actual watched namespace as env var (os.Setenv won't work)	2018-02-06 17:40:06 +01:00
Sergey Dudoladov	8b7bbde06e	Make env var overwrite configmap setting for watching namespaces	2018-02-06 16:12:47 +01:00
Sergey Dudoladov	ea84f9d577	Rename the configmap 'namespace' entry to avoid confusion with the map's owm namespace	2018-02-06 15:09:00 +01:00
Sergey Dudoladov	ec6799f34a	Overwrite scalyr api key if the relevant env variable is present in the operator pod	2018-01-12 14:56:14 +01:00
Oleksii Kliukin	23011bdf9a	Migrate only master pods. Migrate single masters. (#199 ) Avoid migrating replica pods, since they will be handled by the node draining anyway (the PDB specifies that only masters are to be kept). Allow migration of the single-pod clusters.	2018-01-09 11:55:11 +01:00
zerg-junior	bb5ce6cbbe	Merge pull request #195 from zalando-incubator/databases-rest-endpoint Add a REST endpoint to list databases in all clusters	2018-01-09 11:53:32 +01:00
Oleksii Kliukin	8e99518eeb	Improve behavior on node decomissionining (#184 ) * Trigger the node migration on the lack of the readiness label. * Examine the node's readiness status on node add. Make sure we don't miss the not ready node, especially when the operator is killed during the migration.	2018-01-04 11:53:15 +01:00
Oleksii Kliukin	5c8bd04169	Sort database by name.	2017-12-22 15:48:13 +01:00
Oleksii Kliukin	6102b0368c	Merge remote-tracking branch 'origin/databases-rest-endpoint' into databases-rest-endpoint # Conflicts: # pkg/apiserver/apiserver.go # pkg/controller/status.go	2017-12-22 13:08:50 +01:00
Oleksii Kliukin	9720ac1f7e	WIP: Hold the proper locks while examining the list of databases. Introduce a new lock called specMu lock to protect the cluster spec. This lock is held on update and sync, and when retrieving the spec in the API code. There is no need to acquire it for cluster creation and deletion: creation assigns the spec to the cluster before linking it to the controller, and deletion just removes the cluster from the list in the controller, both holding the global clustersMu Lock.	2017-12-22 13:06:11 +01:00
Sergey Dudoladov	b8bf97ab76	Integrate comments from code reviews	2017-12-22 12:53:57 +01:00
Sergey Dudoladov	011458fb05	Add a REST endpoint to list databases in all clusters	2017-12-21 17:28:55 +01:00
zerg-junior	3c178f68df	Warn on infrastructure-roles.yaml format violations (#177 ) Emit a warning if there are unprocessed entries in the infrastructure-roles secret.	2017-12-15 17:21:41 +01:00
Oleksii Kliukin	dd0affc390	Tweak our reaction to the cluster upgrade process. Previously, the operator started to move the pods off the nodes to be decomissioned by watching the eol_node_label value. Every new postgres pod has been created with the anti-affinity to that label, making sure that the pods being moved won't land on another to be decomissioned node. The changes introduce another label that indicates the ready node. The new pod affinity will esnure that the pod is only scheduled to the node marked as ready, discarding the previous anti-affinity. That way the nodes can transition from the pending-decomission to the other statuses (drained, terminating) without having pods suddently scaled to them. In addition, rename the label that triggers the start of the upgrade process to node_eol_label (for consistency with node_readiness_label) and set its default vvalue to lifecycle-status:pending-decomission.	2017-11-30 14:11:49 +01:00
Murat Kabilov	86803406db	use sync methods while updating the cluster	2017-11-03 12:00:43 +01:00
Oleksii Kliukin	eba23279c8	Kube cluster upgrade	2017-10-19 10:49:42 +02:00
Murat Kabilov	6c4cb4e9da	Perform manual failover during the scale down	2017-10-16 17:41:23 +02:00
Murat Kabilov	3b32265258	Set status of the cluster on sync fail/success	2017-10-12 15:10:42 +02:00
Murat Kabilov	8d5faaa5a5	return idle status when worker has nothing to do	2017-10-11 15:42:20 +02:00
Murat Kabilov	83c8d6c419	Extend diagnostic api with worker status info	2017-10-11 12:26:09 +02:00
Murat Kabilov	32aa7270e6	Use round-robin strategy while assigning workers	2017-10-09 16:56:27 +02:00
Murat Kabilov	a35e9c6119	move from tpr to crd	2017-10-06 15:12:08 +02:00
Murat Kabilov	9a66e09b88	cluster history api endpoint	2017-09-26 14:30:45 +02:00
Murat Kabilov	f77852a152	store time of the cluster event	2017-09-26 13:17:23 +02:00
Murat Kabilov	4db5bd13d1	delete cluster key from the clusters list only when delete procedure is finished	2017-09-04 18:48:03 +02:00
Murat Kabilov	899c0bef45	Use warningf instead of warnf	2017-08-30 14:35:56 +02:00
Murat Kabilov	53ceede3cb	show worker queue size in the cluster status	2017-08-28 12:05:33 +02:00
Murat Kabilov	83760ebbef	discard cluster events from the queue on cluster delete; delete cluster from the clusters map before deleting cluster itself	2017-08-17 12:24:23 +02:00
Murat Kabilov	f2c23021bb	generate clusterEvent queue key in a separate function	2017-08-17 12:20:03 +02:00
Murat Kabilov	dad8e2f49f	make cluster event queue consumption non-blocking	2017-08-15 16:03:19 +02:00
Murat Kabilov	82d5583809	add diagnostic api http server	2017-08-15 12:20:09 +02:00
Murat Kabilov	51fdfb90f7	log cluster and controller events in the ringlog via logrus hook	2017-08-15 12:16:09 +02:00
Murat Kabilov	82f58b57d8	add cluster and controller methods for getting status	2017-08-15 12:11:06 +02:00
Murat Kabilov	58572bb43f	move controller config to the spec package	2017-08-15 11:41:46 +02:00
Murat Kabilov	5470f20be4	always pass a cluster name as a logger field	2017-08-15 10:29:18 +02:00
Murat Kabilov	e26db66cb5	start all the log messages with lowercase letters	2017-08-15 10:12:36 +02:00
Murat Kabilov	cf663cb841	Fix golint warnings	2017-08-01 16:08:56 +02:00
Murat Kabilov	c02a740e10	Fix setting debug logger level	2017-08-01 11:51:03 +02:00
Murat Kabilov	6183203f4d	fix cluster event queue processing	2017-07-31 10:30:49 +02:00
Murat Kabilov	2fe22ff614	Remove pod dispatcher	2017-07-27 14:16:49 +02:00
Murat Kabilov	3ad4b127c4	Fix graceful shutdown graceful shutdown of goroutines on operator exit	2017-07-27 12:54:22 +02:00
Murat Kabilov	1f8b37f33d	Make use of kubernetes client-go v4 * client-go v4.0.0-beta0 * remove unnecessary methods for tpr object * rest client: use interface instead of structure pointer * proper names for constants; some clean up for log messages * remove teams api client from controller and make it per cluster	2017-07-25 15:25:17 +02:00
Oleksii Kliukin	4455f1b639	Feature/unit tests (#53 ) - Avoid relying on Clientset structure to call Kubernetes API functions. While Clientset is a convinient "catch-all" abstraction for calling REST API related to different Kubernetes objects, it's impossible to mock. Replacing it wih the kubernetes.Interface would be quite straightforward, but would require an exra level of mocked interfaces, because of the versioning. Instead, a new interface is defined, which contains only the objects we need of the pre-defined versions. - Move KubernetesClient to k8sutil package. - Add more tests.	2017-07-24 16:56:46 +02:00
Oleksii Kliukin	e0dacd0ca9	Remove an unused export.	2017-06-08 16:17:01 +02:00
Murat Kabilov	e104a67260	Fix resync of the clusters	2017-06-08 11:51:48 +02:00
Oleksii Kliukin	bc0e9ab4bc	Add error checks per report from errcheck-ng	2017-06-08 10:41:44 +02:00
Oleksii Kliukin	dc36c4ca12	Implement replicaLoadBalancer boolean flag. (#38 ) The flag adds a replica service with the name cluster_name-repl and a DNS name that defaults to {cluster}-repl.{team}.{hostedzone}. The implementation converted Service field of the cluster into a map with one or two elements and deals with the cases when the new flag is changed on a running cluster (the update and the sync should create or delete the replica service). In order to pick up master and replica service and master endpoint when listing cluster resources. * Update the spec when updating the cluster.	2017-06-07 13:54:17 +02:00
Oleksii Kliukin	7b0ca31bfb	Implements EBS volume resizing #35 . In order to support volumes different from EBS and filesystems other than EXT2/3/4 the respective code parts were implemented as interfaces. Adding the new resize for the volume or the filesystem will require implementing the interface, but no other changes in the cluster code itself. Volume resizing first changes the EBS and the filesystem, and only afterwards is reflected in the Kubernetes "PersistentVolume" object. This is done deliberately to be able to check if the volume needs resizing by peeking at the Size of the PersistentVolume structure. We recheck, nevertheless, in the EBSVolumeResizer, whether the actual EBS volume size doesn't match the spec, since call to the AWS ModifyVolume is counted against the resize limit of once every 6 hours, even for those calls that shouldn't result in an actual resize (i.e. when the size matches the one for the running volume). As a collateral, split the constants into multiple files, move the volume code into a separate file and fix minor issues related to the error reporting.	2017-06-06 13:53:27 +02:00
Murat Kabilov	1fb05212a9	Refactor teams API package	2017-05-30 10:14:30 +02:00
Murat Kabilov	009db16c7c	Use queues for the pod events (#30 )	2017-05-23 15:24:14 +02:00
Murat Kabilov	c470bd6646	reset cluster error on successful update or sync (#29 )	2017-05-22 15:45:38 +02:00
Oleksii Kliukin	bc17897478	Run sync cluster when previous add failed. (#28 )	2017-05-22 15:27:26 +02:00
Oleksii Kliukin	afce38f6f0	Fix error messages (#27 ) Use lowercase for kubernetes objects Use %v instead of %s for errors Start error messages with a lowercase letter.	2017-05-22 14:12:06 +02:00
Murat Kabilov	4acaf27a5d	Remove etcd requests (#25 ) update glide	2017-05-19 17:18:37 +02:00
Murat Kabilov	d34273543e	Fix the golint, gosimple warnings	2017-05-18 17:38:54 +02:00
Murat Kabilov	233e8529c1	Return error instead of logging it	2017-05-18 17:24:44 +02:00
Murat Kabilov	356be8f0f1	skip clusters with invalid spec	2017-05-16 16:46:37 +02:00
Oleksii Kliukin	5adceceb36	go fmt run	2017-05-12 17:48:25 +02:00
Oleksii Kliukin	03064637f1	Allow disabling access to the DB and the Teams API. Command-line options --nodatabaseaccess and --noteamsapi disable all teams api interaction and access to the Postgres database. This is useful for debugging purposes when the operator runs out of cluster (with --outofcluster flag). The same effect can be achieved by setting enable_db_access and/or enable_teams_api to false.	2017-05-12 17:40:48 +02:00
Murat Kabilov	92d7fbf372	replace github.bus.zalan.do with github.cm/zalando-incubator	2017-05-12 11:50:16 +02:00
Murat Kabilov	1b82009151	Command exec inside the Pod method	2017-05-12 11:41:36 +02:00
Murat Kabilov	fd449342e5	Use Kubernetes API instead of API group	2017-05-12 11:41:36 +02:00
Oleksii Kliukin	6983f444ed	Periodically sync roles with the running clusters. (#102 ) The sync adds or alters database roles based on the roles defined in the cluster's TPR, Team API and operator's infrastructure roles. At the moment, roles are not deleted, as it would be dangerous for the robot roles in case TPR is misconfigured. In addition, ALTER ROLE does not remove role options, i.e. SUPERUSER or CREATEROLE, neither it removes role membership: only new options are added and new role membership is granted. So far, options like NOSUPERUSER and NOCREATEROLE won't be handed correctly, when mixed with the non-negative counterparts, also NOLOGIN should be processed correctly. The code assumes that only MD5 passwords are stored in the DB and will likely break with the new SCRAM auth in PostgreSQL 10. On the implementation side, create the new interface to abstract roles merge and creation, move most of the role-based functionality from cluster/pg into the new 'users' module, strip create user code of special cases related to human-based users (moving them to init instead) and fixed the password md5 generator to avoid processing already encrypted passwords. In addition, moved the system roles off the slice containing all other roles in order to avoid extra efforts to avoid creating them. Also, fix a leak in DB connections when the new connection is not considered healthy and discarded without being closed. Initialize the database during the sync phase before syncing users.	2017-05-12 11:41:35 +02:00
Murat Kabilov	2370659c69	Parallel cluster processing Run operations concerning multiple clusters in parallel. Each cluster gets its own worker in order to create, update, sync or delete clusters. Each worker acquires the lock on a cluster. Subsequent operations on the same cluster have to wait until the current one finishes. There is a pool of parallel workers, configurable with the `workers` parameter in the configmap and set by default to 4. The cluster-related tasks are assigned to the workers based on a cluster name: the tasks for the same cluster will be always assigned to the same worker. There is no blocking between workers, although there is a chance that a single worker will become a bottleneck if too many clusters are assigned to it; therefore, for large-scale deployments it might be necessary to bump up workers from the default value.	2017-05-12 11:41:35 +02:00
Murat Kabilov	a7c57874d5	Do not create roles if cluster is masterless fix pod deletion	2017-05-12 11:41:34 +02:00
Murat Kabilov	da438aab3a	Use ConfigMap to store operator's config	2017-05-12 11:41:34 +02:00
Murat Kabilov	08c0e3b6dd	Use unified type for the namespaced object names	2017-05-12 11:41:34 +02:00
Oleksii Kliukin	71b93b4cc2	Feature/infrastructure roles (#91 ) * Add infrastructure roles configured globally. Those are the roles defined in the operator itself. The operator's configuration refers to the secret containing role names, passwords and membership information. While they are referred to as roles, in reality those are users. In addition, improve the regex to filter out invalid users and make sure user secret names are compatible with DNS name spec. Add an example manifest for the infrastructure roles.	2017-05-12 11:41:33 +02:00
Murat Kabilov	db53134cbd	Skip syncing Pods	2017-05-12 11:41:33 +02:00
Murat Kabilov	101dc06acb	Better logging for teams api calls	2017-05-12 11:41:32 +02:00
Murat Kabilov	bb4fec25ae	Fix deletion of the failed cluster; more debug messages	2017-05-12 11:41:32 +02:00
Murat Kabilov	ce90a54cf9	create key in the cluster map on cluster creation failure	2017-05-12 11:41:32 +02:00
Murat Kabilov	852c5beae5	Check etcd key availability for the new cluster	2017-05-12 11:41:31 +02:00
Oleksii Kliukin	8268b07ad2	Set logger level per package instead of doing this globally	2017-05-12 11:41:30 +02:00
Oleksii Kliukin	3a4c6268be	Increase log verbosity, namely for object updates. - add a new environment variable for triggering debug log level - show both new, old object and diff during syncs and updates - use pretty package to pretty-print go structures -	2017-05-12 11:41:29 +02:00
Murat Kabilov	c2d2a67ad5	Get config from environment variables; ignore pg major version change; get rid of resources package;	2017-05-12 11:41:29 +02:00
Murat Kabilov	79a6726d4d	Increase logging verbosity, restructure code	2017-05-12 11:41:28 +02:00
Oleksii Kliukin	48ba6adf8a	Avoid calling Team API with an expired token. Previously, the controller fetched the Oauth token once at start, so eventually the token would expire and the operator could not create new users. This commit makes the operator fetch the token before each call to the Teams API.	2017-05-12 11:41:28 +02:00
Murat Kabilov	6f7399b36f	Sync clusters states * move statefulset creation from cluster spec to the separate function * sync cluster state with desired state; * move out from arrays for cluster resources; * recreate pods instead of deleting them in case of statefulset change * check for master while creating cluster/updating pods * simplify retryutil * list pvc while listing resources * name kubernetes resources with capital letter * do rolling update in case of env variables change	2017-05-12 11:41:27 +02:00
Murat Kabilov	486c8ecb07	use neutral name for set cluster status function	2017-05-12 11:41:26 +02:00
Murat Kabilov	1c6e7ac2e7	loadBalancerSourceRanges update	2017-05-12 11:41:26 +02:00
Murat Kabilov	fc127069ab	remove unnecessary ControllerNamespace	2017-05-12 11:41:26 +02:00
Murat Kabilov	416dace289	get rid of arrays in the kuberesources; use shorter form of checking for errors	2017-05-12 11:41:26 +02:00
Murat Kabilov	ae77fa15e8	Pod Rolling update introduce Pod events channel; add parsing of the MaintenanceWindows section; skip deleting Etcd key on cluster delete; use external etcd host; watch for tpr/pods in the namespace of the operator pod only;	2017-05-12 11:41:25 +02:00
Murat Kabilov	dfde075c66	Use TPR object namespace while creating its objects	2017-05-12 11:37:09 +02:00
Murat Kabilov	6e2d64bd50	Create human users from teams api	2017-05-12 11:37:09 +02:00
Murat Kabilov	58506634c4	Create pg users	2017-05-12 11:37:09 +02:00
Murat Kabilov	abb1173035	Code refactor	2017-05-12 11:37:09 +02:00
Murat Kabilov	75e6bfa55c	makefile improvements	2017-05-12 11:37:07 +02:00
Oleksii Kliukin	e96f8a80ee	Option to run the operator out of cluster.	2017-05-08 12:10:27 +02:00
Oleksii Kliukin	b3a9516bae	Add a missing file.	2017-05-08 12:10:26 +02:00
Oleksii Kliukin	e5e0e3a148	Use camelCase.	2017-05-08 12:10:26 +02:00
Oleksii Kliukin	38bc9da25a	WIP: allow operator to run both in- and out- of cluster.	2017-05-08 12:10:26 +02:00
Murat Kabilov	5b5a64e55d	Check if etcd service has its port exposed	2017-05-08 12:10:26 +02:00
Murat Kabilov	d5a7683a38	some refactoring	2017-05-08 12:10:26 +02:00
Murat Kabilov	256ff37c19	refactor file tree structure	2017-05-08 12:10:25 +02:00

... 2 3 4 5 6

265 Commits