
The Database: The Elephant in the Room
Om avsnittet
Why the Database Is Different
Grant Fritchey, who calls himself the Scary DBA and has been an enterprise DBA for about 20 years, and Jonathan Hickford, a product manager, both from Redgate, explain why databases are the elephant in the room of continuous delivery. Grant's answer is data persistence. Releasing software is straightforward: "You take an old piece of software, you throw it away, put the new piece of software in, you're done." You could do the same with a database, "but then the phone starts ringing because the business is freaking out because there's no more data left in the system."
Jonathan adds that data comes from several places, some generated by users and some configuration data from development, and often flows in the opposite direction from the code in a pipeline. A database may also be read by several applications and BI jobs, so a change has to be safe for all of them. Grant says data migrations for something like a new non-null column run against a live system and can cause downtime, and that part of the push for NoSQL comes from the difficulty of database deployments.
Reviewing the Script the Day Before Is Too Late
Grant hates the common DBA habit of looking at the script the day before it goes to production, which means the first run happens in production. Matty says he's seen DBAs who ran the developers' scripts without reading them, which he says adds no value. Grant, throwing rocks at himself, says he learned the hard way that on Thursday you can't stop a Friday deployment: "What you have to stop is bad development," which means going to the scrums and standups and knowing what's happening six months, six weeks or six days before, not six hours.
Jonathan says the core of continuous delivery is fast feedback, so a data problem should be found in development, and production deployment "should be boring." Grant says he'll steal that: people always ask how your plane flight was, and you don't want an exciting one, "because you don't want an exciting production deployment."
Tooling and Source Control
Grant says step one, though it sounds bad from a tool vendor, is tooling, since you can't get a database under source control by hand. You need differential scripts, alter scripts and migration scripts that move data, not create scripts rerun over and over. His objection to ORM tools is that they only produce 1.0-style code, such as create table, when existing data has to persist. Jonathan says once the database is in source control you can repeat things, and test a refactoring against realistic data to find out it would take three hours over millions of rows. Grant adds that even knowing it takes three hours lets you schedule for it.
The Myth of Rollback
Matty says knowing you can't go back and are always going forward is a reason to test deployments continuously, since the old habit of testing your backout never happened anyway. Grant says only two rollbacks really work: restoring or undoing a snapshot right after the release, or redeploying forward. A script to roll back pieces is "such a lie because the data comes in over time and ain't nobody wants to get rid of it." So the deployment process itself has to fix mistakes.
He says test environments should be as close to production as possible with cleaned data, for example without emails or personal details, and that sometimes it's fine to refresh QA when it burns down. Matty describes a former 3.5 terabyte data warehouse multiplied across more than 20 environments and being told there was no time to make test data. Jonathan says you don't need every environment to be a copy: one for data size, one for data complexity, and lighter ones for fast changes. He suggests testing sharding and replication earlier, and prefers "the short, fat pipeline rather than the long, thin pipeline."
Blue-Green for a Database?
A question from IRC asks how to blue-green a database change. Jonathan says it's possible with patterns, such as splitting a column by letting the application or data access layer handle both versions and migrating slowly at quiet times. Trevor asks about schema changes and Grant's answer is one word: badly. He's done it with triggers keeping two copies in sync, and calls it a "massive undertaking and fraught with horrific danger," like Indiana Jones. Jonathan says if a DBA says the database isn't the place to tackle this, they're probably right, and offers versioned schemas as an alternative.
Matty adds that in agile shops the throwaway shim, like a bridge you won't need once the stream is gone, is a hard sell to a product owner with a user story waiting.
How DBAs Feel About DevOps
Grant says it's a mixed bag. The camp that upsets him most says everything's fine because they've always done it that way, and Matty's reply is "We have always been at war with East Asia." To reach them you have to document their pain: how long a deployment took, whether there was downtime or data loss, how long recovery took, and the same for the developers who hop through hoops before production, and take that to management. Others are ready, and Jonathan says at DevOpsDays and other conferences someone always asks in the Q&A "what are you doing about the data?"
Grant says he's a "completely lazy bastard" who takes advantage of what development teams have already worked out about source control, labeling and branching: "it's down to more labor than thought." He supported 10 development teams by automating builds and deployments, letting developers write their own T-SQL and reviewing for 15 to 20 minutes a day per team. Jonathan says the best visits he's made had a DBA, a developer and a sysadmin in the room.
Common Tooling, Testing, and Meeting in the Middle
Matty worries about silos created by tools, like data developers stuck on a different source control system. Jonathan says vendors are becoming more agnostic and that using one system lets you make atomic commits containing both database and application changes, which he calls probably a prerequisite for continuous delivery. Grant says without everything versioned together, people had to ask which database version went with which application.
On testing, Jonathan mentions tSQLt and DBUnit, and Grant says you don't write a test for a table. What matters are tests on destructive changes, that a migration moved the same number of rows and the updated data looks as expected, for confidence and boring deployments. Jonathan adds monitoring after release. Grant says DBAs need to learn tools like TeamCity, Jenkins, Octopus and PowerShell, because "we can only automate all of that. So what is your job again?" and they need to meet developers in the middle, not stand "athwart the bridge stopping you from going to production." Matty says sysadmins need to get smarter about data too.
Fler avsnitt
Visa alla avsnitt av Arrested DevOpsArrested DevOps med Matt Stratton, Trevor Hess, Jessica Kerr, and Bridget Kromhout finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.