Showing posts with label mission critical. Show all posts
Showing posts with label mission critical. Show all posts

Tuesday, October 12, 2021

Managing Mission-critical Products in Flight (Literally!) with Ben Berry (#PNSQC2021 Live Blog)

 


Headshot of Ben Berry
This is an interesting area I had not considered.  Ben Berry is the CEO of AirShip Technologies Group. Their product is unmanned aerial vehicles. think drones that can reach heights to release and deploy satellites. 

These are the definition of mission-critical products. If they don't perform as required, catastrophe can result, either in space, on the ground, or anywhere in between. To say the software used in these capabilities is complex and has some very specific and demanding requirements is an understatement. 

To quote Ben:


"AirShip Technologies Group’s VX Unmanned Aerial System is a reusable air platform for high altitude, micro-rocket launch for payload placement of 5G communications in low earth orbit (LEO) of up to six nanosatellite (6.5 lbs. each) or 100 picosatellite (2.2 lbs. each), or 1,000 femtosatellite (0.22 lbs. each). The autonomous VX delivers communications technologies to improve Satellite Communications (SATCOM) link resilience, throughput, and reduced user equipment when compared to SpaceX."

Okay, wow... now that got my attention. I admit the thought of testing something like this would be both amazing and terrifying.



Let's think about how we might test the following:

- R&D focuses on scientific benefits and commercial applications of on-demand microSAT 5G communications deployment; 
- just-in-time launch capabilities for expanded 5G communications via miniaturized LEO satellites. 
- Design objectives include resilient, interconnected 5G mesh communications; 
- exploitation of ultimate high ground of space communications; 
- space empowered SATCOM link resilience; 
- strategic space force projection and operational agility; 
- communications via 100x factor 5G bandwidth; 
- Modular interoperability among microSATs for communications that meet growing bandwidth and resiliency requirements.


Mind blown? Yeah, so is mine!



Any time I tend to think I've had a chance to work on some complex projects, I see something like this and go "nope, noooope". This sounds incredibly daunting and yet there is a part of me that thinks testing something like this would be a total rush (LOL!). However, to quote Dirty Harry... "A [person] has to know their limitations." 

Tuesday, October 13, 2020

PNSQC 2020 Live Blog: Quality Focused Software Testing In Critical Infrastructure with Zoë Oens




We are all familiar with the "Iron Triangle" where we get three sides; Faster, Cheaper, Better... Pick Two, because that's all you will be able to achieve.  One hundred percent coverage is unachievable. Especially if the goal is to save money in the process. still, what is the answer when you have to test software on a  critical system? when I say critical, and when Zoë says critical, we are talking things like the power systems, water, medical, banking.  Bugs in production are not trivial.



So what do we do when it comes to testing CI (critical infrastructure) software? How close to one hundred percent coverage can be achieved? If a feature fails, what is the fallout? How many are affected? While not specifically software related, some of you may know/remember that I have direct experience with a Critical Infrastructure failure in my community. My town of San Bruno suffered a major gas line explosion in 2009. Maly lives lost and many homes destroyed. It took years to rebuild and in many ways, the scars and memories are still fresh. 

When we are looking at testing in these critical areas, we have to be able to prioritize and determine the mission-critical stuff. Granted, we can't just declare everything as mission-critical but dealing with the electricity grip or gas supply, yes, critical becomes more meaningful. 

Zoë mentioned her time in manufacturing helped her approach the questions and issues where CI comes into play. Make sure that there is a spread of knowledge. CI is an area where there should be no silos. No one person should be the one to know how to fix problems when they occur, both at a coding and an operational level.

An emphasis on test writing, test setup, and test execution needs to be codified and well understood, as well as run frequently and aggressively to be sure that as many cases have been focused on and addressed as possible. Automation falls into this spare especially if the steps are codified and well known. anything repetitive, especially if it is rote repetitive, should be automated. The goal in CI environments is to be able to run testing consistently and effectively. While there is an upfront cost here, this cost can be amortized over time, especially if any of the tests are long-lived.

CI environments can sound scary but the key takeaway to me is to be mindful and diligent, as well as look for repetitive areas that can be codified, confirmed, and repeated.

Tuesday, January 21, 2014

Questioning My Expertise

It's been while since I've posted anything. Again, that's both on purpose and not. I felt the need to disconnect for a bit and take care of some other things in and around my life that have, frankly, suffered a bit from neglect (my home, my yard, some much needed family time, and a project that is going to be the main part of this entry today).

Back in November, I posted about a "tank crash" that I suffered, one in which a fish tank that, in some way, shape or form I'd been running, mostly non stop, for close two decades came to an ignominious end. From a full community to zero survivors, and from zero survivors to a tenuous hold on  new community and eco-system. In this process, I have had to come face to face with the fact that everything I thought I knew, and all of the methods and techniques that just "worked" for me basically just stopped working.

What does a complete eco-system crash entail? What does one do when an entire world they were maintaining comes crumbling down? Yeah, this may be a bit over-dramatic, but trust me, I've prided myself on keeping fish alive and thriving for over a decade. To have them all die in a short period of time, even after taking major precautions and performing heavy and expensive interventions, has been very frustrating. The situation reached a point where a "baptism by fire" were necessary. Well, OK, not really, but a cleansing of Clorox, high heat and desiccation was very much involved. For those who have ever wondered what a complete purge of a system involves, it basically goes something like this:

- deconstruct everything in the tank. That includes removing all filter media, all rock work, any real or plastic plants, any decorative hiding spaces, and all of the tanks substrate (in my case, ranging from fine sand to pea sized gravel.

- drain all of the water: all 65 gallons worth. Needless to say, the front and back yard plants and trees yard received a lot of attention that day.

- take all of the filter media out of the various filters (sponges, ceramic tubes, plastic inlet/outlet tubes, carbon bags, phosphate absorbers, etc.) and put them in the dishwasher (having run it empty with no soap prior to this for preparation. the dishwasher, acted as a pseudo-autoclave for sterilization purposes).

- take all of the decorative rock work, filter media baskets, decorative materials, etc. and perform the same process in the dishwasher

- gather all sand and gravel into a bucket and boil in various pots until everything had been heated for at least 15 minutes in boiling water.

- soak the boiled rock in a 6.25% Chlorine bleach solution (that's a cup of bleach to a gallon of hot water) for 72 hours, then rinse with clean water and let soak in water treated with tap water conditioner (to remove any remaining chlorine).

-wipe down the tank, inside and out, with the same 6.25% solution and let dry. Rinse and fill with fresh water and tap water conditioner.

- put everything back into the tank and test over several days to make sure any chlorine or other trace metals were all gone.

- add a biological filter agent to the tank and introduce a "cycling population". This may sound cruel, but it needs to be done, and a group of fish are the best to do this. For my purposes, a school of six Giant Danios were my literal "canaries in the coal mine". A month later, and they have been doing fine, as have four Boesmani Rainbows I've introduced since the tank has been cycled.

Sounds like everything is working great, huh? Well, not quite. This morning, two fish I recently purchased and was taking care of with the hope of introducing into the main tank after a quarantine showed many of the same hallmark symptoms of the disease I had just tried so hard to eradicate. WHY?!! What is going on here? Why am I seeing such large scale infestations when I wasn't seeing them before? Why was it happening again? I'd been using a quarantine tank. I'd been changing the water. I'd been feeding in very small amounts and monitoring them. All of the things that I had figured would be to their benefit, and yet, this morning, I found two dead red severums in my quarantine tank. AGGGHHHHH!!!!

Fortunately, after years of being a tester, I decided to stop cursing my bad luck and start thinking systematically about what could be the problem. First, my quarantine tank. It's six gallons. Not large, but then it's not really meant to be. It's only going to hold one or two fish at any given time, and then just for brief periods. The fish that inhabit the tank are all juvenies, and the tank is outfitted with filtration, aeration, light, heat and all the things necessary to keep them healthy, or so I thought. The rainbows had been through the quarantine process, and had suffered no ill effects. Why did these fish have such a different fate?

Part of it has to do with a different morphology. Rainbowfish are cyprinids, and as such they are highly active, but very efficient fish. They don't eat a lot, and they produce small amounts of waste compared to the red severums. Though the classic yarn of "an inch of fish per gallon" was maintained, the four rainbows produced lots less waste than the two similarly sized red severums. I had figured regular water changes, to the tune of a gallon every other day, would be sufficient to maintain good water quality. It made sense for the rainbows, and may have been overkill. For the severums, frankly, it may have been too little.

There was one other little piece of the puzzle that was introduced, and frankly, I hadn't even realized it. Since the small half bathroom has a sink and storage for everything I use in the aquarium hobby, it just made the most sense to set up the quarantine tank right there in the bathroom. Of course, there's also other activities that take place in a bathroom, and to help keep said room "pleasant", we'd plugged in a wall vaporizing "air freshener, the kind that heats up essential oils to make a pleasant smell for the small room. Having used it for so long, I hadn't even really given it much of a thought, except to notice that, a couple of days ago, there was a stronger smell. The reason? Christina had changed the small bulb and put in a new scent, one that was more noticeable. For people, not such a big deal. For the respiration of fish? Think of how it feels to breathe in super concentrated pine oil when you clean something. Now imagine you can't escape it. Yep, dare I say it, that might have had something to do with it. Were the fish in a more open room, or had there been stronger venting of the room, then it wouldn't have been an issue, but in such an enclosed space, I'm guessing that exposure could prove to be lethal.

So here I am, an empty quarantine tank, two casualties, and a bit of frustration. However, I can take some solace in one aspect of this... the system worked as it was designed. While my fish looked to have succumbed to a common ailment (one that, realistically speaking, all fish have) the environment I had unwitting set up, for the best of purposes, caused the condition to manifest, and to become fatal to them. On the positive side, if there can truly be one, is that I didn't release these fish into the main tank, where the infection they carried would have spread to the other fish and started the whole process over again.

It's easy to start second guessing yourself when something that works well for so long breaks down. When that happens, it seems like everything you do from that point on no longer yields success.  It's enough to make one want to throw in the towel completely, but I also realize that systems, especially ecological ones, are complex. It takes time to reach a new equilibrium, and to get new communities to thrive again. There will invariably be mis-steps, second guessing, and a loss of overall confidence in one's efforts. It's especially frustrating when lives are on the line, even if they may just be "fish that only cost a few bucks".  I feel it acutely every time an animal dies in my care, because it forces me to think and ask "what could I have done differently? How could I have been better prepared for this?" Additionally, having a space that I can use to isolate and care for  individuals in a manner that will help them thrive, as well as firewall them from others should something go wrong when we bring them home, is now more important than ever. That first quarantine is vital. Making that environment as healthy as humanly possible is critical, and little things that we often don't think about can have huge repercussions. Here's hoping I've learned enough these past few weeks that future additions will be able to live and thrive, if not disease free, then with as little chance of having those issues coming to the surface as humanly possible.

Thursday, September 5, 2013

When a Wish Falls into Your Lap

I cannot say the word "PumpKing" without this image
popping into my head ;).
This is a bit amusing, a bit serendipitous, and a whole lot of awesome all rolled into one. I don't typically talk about my company in this blog unless it's to make a general point or to highlight something I find interesting, and today's comments fall squarely into the latter.


For quite awhile now,  I've been hoping for a window of opportunity to look at and open up some venues of more technical engagement for both myself and the software testing team, in ways that go beyond writing automation scripts. One of those opportunities literally fell into our laps a couple of weeks ago.


My company has a long running and enterprise level product that was originally written in Perl. As part of the development process, over the years, we have had a movable role of build manager, code health manager, deployment specialist, front line troubleshooter, and whip cracker all rolled into one. The term (that has a lot of history in the Perl community) for this role is the "PumpKing" (holder of the pumpkin, they who keeps the system afloat, the puller of the strings, themz who pulls the taps to keep the good stuff flowing, call it what you will).

The responsibility of PumpKing has been handed off week after week to the various programmers in a round-robin fashion. A few weeks ago, I made an aside about one of the programmers (who I happen to be good friends with) not living up to their responsibility as "PumpKing". This had to do with running stand-up that morning, i.e. it starting late, and my comment was totally in jest. Their response was (and with a smile, I might add):


"Well, Michael, if you would like to get into the PumpKing rotation, we can certainly arrange that!"


I chuckled, thinking the comment was likewise in jest.


It wasn't.


At our next Engineering team meeting, said programmer made a motion that "all software testers should join the PumpKing rotation". I snickered again, thinking this was follow-up on the joke. What I was not prepared for was when our VP of Engineering said "I think that's a great idea!"


After I did my double take, and realized that I was not being punked, I stopped and thought about what an awesome opportunity had just been handed to us. Why is this an awesome opportunity? Because in one fell swoop, the potential number of "active duty build and deployment engineers" just went up considerably. For a group of software testers looking for an excuse/opportunity for more technical engagement with the development team, we were just given the keys to a gold mine.


Needless to say, there's a lot of learning to be done so that we can do end to end builds, merges, CI, testing, deployment to various environments and interacting with Ops to make the final pushes to production. Additionally, the PumpKing runs stand-up during their week, is on call for issues, interacts with front line support for rapid response of issues, etc. It's kind of a big deal!


I'm excited because this is a real opportunity for those of us that want to have an avenue to better understanding of the technical underpinnings to really get into it. We have the potential to do hot fixes if needed, we are responsible for pushes in our development and staging environments, as well as with pushes to production. It's visible, it has the potential to get wild, and it's a great way to get knee deep in the muck and grime of the real code base and understand how everything fits together. In short, it's a dream come true for a software tester wishing there were some way to more effectively blur the lines between programmer and tester.


These past couple of weeks have been all about learning the ropes, practicing along with the current PumpKing, and getting ready for my turn in the saddle... which starts this coming Monday morning and extends to the following Monday morning.


Let the festivities begin ;).

Monday, July 8, 2013

Cascading Fail: A Crash That Hits Close to Home

As many of you are aware, there was an airliner coming from Seoul, Korea that crashed upon landing at San Francisco International Airport Saturday, July 6, 2013. Much has been said in the news about the crash, and much will probably be said in the following days and weeks. My point for this post is not  to talk about the crash, or the myriad of systems that were or were not available. It doesn't change the fact that an airline crashed, scores of people were injured, two girls were killed, and a major International Airport was brought to a grinding halt while the events that followed played out.

This crash had a direct impact on me, though, in two ways. First, it is the cause of my daughter, who has spent the last eight days in Japan, being unable to come home yet. Because of the shut down of the airport, many flights had to be diverted to smaller airports, which were quickly overwhelmed by the increased load. Many flights that were scheduled were cancelled, my daughter's included. Thankfully, due to some herculean and determined efforts by the team of chaperones, as well as the good will and kindness of the city of Narita, Japan, they were able to weather the hiccup. Note, this "hiccup" meant that their return was delayed by three days, including, at the current time, a rerouting through Seattle and an almost 24 hour layover.

The second way that it impacted me is knowing that, for the two girls who were confirmed dead, both were students coming over on an exchange trip from China. Both were mid teenagers. In other words, both were mirroring my own daughter's trip. I am a realist, and I understand Black Swan events, and the likelihood of a repeat performance is way less than a million to one, but that doesn't calm a father who now has new uncertainties and anxieties. Needless to say, it was a little too close to home.

Note, I'm not blaming anyone for this, but in the tester's world that I inhabit, rarely is there such a thing as a "isolated problem". Usually, when something goes wrong, it affects entire systems, and those systems also affect entire systems. The net results of an error could cascade out and cause devastating problems, and the after effects not seen until well after this issue has occurred. Yes, I am one of those people who can see a testing story in everything, but today I am doubly reminded of how early mistakes not caught can ripple out, and we really have no way of knowing just how far the problems  can cascade. We may have a momentary interest, as long as the issue doesn't affect us. Once it does, though, we start to see a much broader world of issues and problems. A good reminder to me in my workaday world to try to find problems early, while the course correction options are much wider and still possible.

Friday, March 16, 2012

Being a Fly on the Wall, Or a Plane

Today was a chance to take a break, get some travel out of the way and get myself from San Francisco up to Calgary for the Calgary Perspectives on Software Testing Workshop (POST).

The fun started when I heard news that United and Continental were merging their computer systems this week. I figured "this should be interesting". It definitely became even more interesting when I discovered, after standing in Air Canada's line for a half hour early this morning, that the flight was being serviced by United Express. Meaning I was in the wrong terminal, and had better get a move on over to the Domestic terminal if I wanted to catch my flight.

I hustled on down, and tried to check in at the kiosk, only for the kiosk to tell me my ticket was out of order and that an agent would have to help me. Turns out my order was placed for a paper ticket, even though I processed the whole thing online and had the receipt to prove I was supposed to be getting an e-ticket.

This comedy of errors was resolved and I made my way through security (actually much faster now that I know that stainless steel plates don't set of metal detectors). After making my way to the counter, showing my passport, waiting just a few minutes and then getting on the plane, I figured this was all over... well, not quite.

It seems setting up and merging the two companies (Continental and United) resulted in a rather big amount of software conversion... conversion with limitations. One of the casualties of this conversion was the ability to electronically match the flight manifest with what was stored in the computer. The Flight crew had to do a manual tally of all passengers, and write down and submit this manifest. Not such a big deal on the surface, except that it delayed the flight by about 30 minutes (not a problem for me personally, but I'm sure those making connections in Calgary were none too thrilled).

I'm sure that we'll see something written up about this in the papers at some point, and it will be labeled as a "glitch". Well, no, it's a little more than that; it's a failure of the system's to integrate and there's little in the way of information being shared with the passengers to describe what is happening. I don't mind the errors. Being silent about it is what irritates me.

Still, all's well that ends well. I made it to YYC, a cab took me to my hotel, and I'm just geeking out and focusing on some work until friends come and pick me up for dinner. Looking forward to my time here in Alberta, thanks go out to those who invited me to be here :).

Saturday, October 8, 2011

BOOK CLUB: How to Reduce the Cost of Software Testing (10/21)

For almost a year now, those who follow this blog have heard me talk about *THE BOOK*. When it will be ready, when it will be available, and who worked on it? This book is special, in that it is an anthology. Each essay could be read by itself, or it could be read in the context of the rest of the book. As a contributor, I think it's a great title and a timely one. The point is, I'm already excited about the book, and I'm excited about the premise and the way it all came together. But outside of all that... what does the book say?

Over the next few weeks, I hope I'll be able to answer that, and to do so I'm going back to the BOOK CLUB format I used last year for "How We Test Software at Microsoft". Note, I'm not going to do a full synopsis of each chapter in depth (hey, that's what the book is for ;) ), but I will give my thoughts as relates to each chapter and area. Each individual chapter will be given its own space and entry.

We are now into Section 2, which is sub-titled "What Should We Do?". As you might guess, the book's topic mix makes a change here. We're less talking about the real but sometimes hard to pin down notions of cost, value, economics, opportunity, time and cost factors. We have defined the problem. Now we are talking about what we can do about it. This part covers Chapter 9.

Chapter 9: Postpone Costs to Next Release by Jeroen Rosink

So how many of us have been on projects where it seems the feature set or the deliverables are really overloaded or far reaching? Is there anything wrong with being ambitious? Of course not. Is there anything wrong with trying to outdo the competition? Again, no, but ask yourself, is everything that is being promised really as important as people are making it out to be? Do we really need to deliver every single one of the features that are listed right now? In many of the projects I have worked on, the answer has proven to be "no" more often than "yes". My guess is you've felt the same way. What if we culled the load a little bit? How about if we really focused on the features that actually mattered? What if we could come to a consensus on what those features were, and deliver those during our window, and then allow some of the other issues/challenges/opportunities to be addressed for the next release window?

Some might say "ah, but that's being defeatist!" Is it? We all know that we do this anyway; there's almost always issues left on the table when a product releases, but we tend to leave them on the table after pulling our hair out trying to code/test/recode/retest and ultimately giving up because if we keep at it, we really will miss our window of opportunity. That happens all the time, and that process is very expensive. Jeroen is suggesting instead that we take a more pro-active approach and rather than reach the "leave things on the table" state in desperation at the end, that we focus on determining what is really important first, and try to avoid that frustrating phase at the end altogether if we can.

Jeroen describes the often challenging tug-of war between postponing features and issues and demanding solutions be found to known issues. Often the challenges of solving a problem "right now" leads to false starts, dead ends and unsatisfactory answers, where putting more time and research into a solution can yield better results. There is no question that some products are born in a cauldron of intense pressure, the time limit the crucible that can bring to the fore a make or break idea that at its heart is pure genius. However, the success rate of that approach is far less than you might imagine. What usually happens is that we have a rushed solution that is often missing a lot in its implementation, and the odds of large scale problems surfacing in the field go way up. Taking time to adequately study a situation, look at potential permutations, and code a solution that answers them usually takes a significant chunk of time, and doing so at a measured pace may well prove to be more effective both in the way of the solution provided and the costs to produce it.

Structuring testing in phases can help identify critical areas and time needed to test adequately and effectively. Activities such as collecting documentation, defining test cases, creating test data, and executing tests all falls into these phases. The most common (and repetitive) activities are specification, preparation and execution.

The primary goal of any software delivery is business value. How often is there a claim from business that everything is important and that everything must be delivered? Under these conditions everything would merit being tested at the same depth. It's true that some mission critical processes may need extended or extraordinary testing. In truth, if all of the features were truly evaluated, we'd see that they will fall under different levels of the MoSCoW principle ("Must have", "Should have", "Could have", "Would like to have"). When presented with the "everything is important" argument, we need to answer back "OK, are you wiling to open us up to the risks of treating everything at the same level of importance?"

Product risks are those that are directly related to a particular product under development. Project risks are those related to the resources and or the costs of particular project. MoSCoW helps classify different risks. This allows a realistic overview where decisions can be made. Late delivery of functionality to meet requirements results in later-stage project risks. After risk-assessment is performed a test strategy can be defined. Not everything needs to be tested to the same extent. Issues discovered can help inform the development or re-evaluation of the testing strategy over time.

It's important to remember that even simple changes can have big impact on a project, so more complex changes will have even more of an impact. An important question to always ask is "What impact will there be if we discover an issue later in the project?"  When issues are discovered, we need to consider their technical impact versus their business value. Does delivering with known issues produce a net positive business gain? Is it worth the risk to try to fix it?

The value of an issue changes during the life of a project. Early on, all issues feel equally important. Later on issue values may need to be examined in context, and determine how valuable fixing that issue at this time really is. Is it important enough to potentially delay release? Costs definitely increase towards the end of a project, especially if a resource or piece of functionality is a must have and it's not fit to be delivered close to the end of the testing cycle. By considering the costs per test phase, it may make sense to postpone some of the solution to a later release.

Monday, October 3, 2011

BOOK CLUB: How to Reduce the Cost of Software Testing (5/21)

For almost a year now, those who follow this blog have heard me talk about *THE BOOK*. When it will be ready, when it will be available, and who worked on it? This book is special, in that it is an anthology. Each essay could be read by itself, or it could be read in the context of the rest of the book. As a contributor, I think it's a great title and a timely one. The point is, I'm already excited about the book, and I'm excited about the premise and the way it all came together. But outside of all that... what does the book say?

Over the next few weeks, I hope I'll be able to answer that, and to do so I'm going back to the BOOK CLUB format I used last year for "How We Test Software at Microsoft". Note, I'm not going to do a full synopsis of each chapter in depth (hey, that's what the book is for ;) ), but I will give my thoughts as relates to each chapter and area. Each individual chapter will be given its own space and entry. Today's entry deals with Chapter 4.

Chapter 4: Opportunity Cost of Testing by Catherine Powell

Catherine starts out the chapter with something we all know but rarely directly address… testing is filled with roads that are never taken. Any time we perform a test, we are deliberating not doing hundreds or thousands of other tests. Much of this is based in practicality. There is no way to ever do “complete testing”, even with simple programs. These “roads not taken” do have a cost though. They are opportunities that we do not explore, and therefore they are avenues that could become issues later. We somewhat vaguely understand that. This chapter puts that idea in to clearer focus. Opportunity costs are what we might gain if we travel down a different road. I may make a certain amount of money at a job that I work. Could I make more money at another job? It’s possible. It’s also possible the company could fold up and I could make a lot less money, or for a time, no money. Still, the opportunity cost is there.

As we test, we also have choices that are often based on time requirements. Do we perform functionality tests or stress tests if we have time to run just one? If we run functionality tests, we incur the opportunity cost of not running stress tests, and vice-versa. It’s a trade-off, and one we have to make at times, but now we can understand that there is indeed a cost for that choice.

Many testing theories work great in a lab or in a hypothetical situation, but there are real world constraints in which all projects operate. There isn’t an unlimited amount of time, or resources, or equipment, so all of those become constraints with a ship window coming due. Because of that, there are opportunities that by necessity must be left on the table, and then those opportunity costs are incurred. So we can see that opportunity costs are genuine. That’s great that we understand that, but how can we actually apply this information effectively? Put simply, there needs to be a priority placed on certain tests, but more than jut a priority, a clear understanding as to why those tests are a priority. Installation and upgrade testing may be given top priority because, if those tests fail, then the software will not work at all. That outweighs the opportunity costs of discovering how many browsers will be able to fully support the application. So choices are made, and understanding the opportunity costs helps to put into perspective what those priorities are. It’s not so much what you are willing to have, it’s what you are mindfully, and with deliberation, willing to give up so that the most important areas do get covered.

To help us with this process Catherine shows us the idea of a decision table, which shows a few details to help us determine what we will be testing and what their value would be. By doing this, we can examine each of the test areas and we can order them by priority and, hypothetically, by direct and opportunity cost. Each item has similar criteria. There’s the actual task, the time it will take to do it, the effects of not doing it, and the priority of the item. I list the priority last because you need to have a clear understanding of the other criteria to make judgment on what the priority actually is.

Opportunity costs are not set in stone. Many details can be changed. The most common variable that changes opportunity costs is time. If something happens that extends the time, then opportunity costs drop, because with more time, more tests can be performed. Of course, this works the opposite way, too. When time estimates need to be cut, then opportunity costs dramatically rise, and the entire decision table may need to be totally reworked. What was once seen as an acceptable trade off of priorities may now no longer be so, and the priorities may need to be reordered or completely revised. We could conceivably add more people to the project, but there’s a cost there, too. If we have enough time to ramp them up to speed, then it’s not too bad a trade off. If we are days before shipping, adding a new person means we have to train them in what to do, which brings a whole new set of opportunity costs regarding the tests that are not being performed while we train the tester.

Ultimately, everything we do has a cost, not just in terms of direct action, but in terms of the things we do not do. There’s always more we could do, and there’s never enough time to do it all. There will always be opportunity costs in testing. We cannot isolate ourselves from them, but we can understand them, and if we follow Catherine’s advice, we can leverage them to our advantage.

Saturday, August 27, 2011

How Fragile is Your System?

My kids had a somewhat sad, but I think important, lesson taught to them this week. We have been involved in a process since March with our main fish tank in our house; we have a 65 gallon tank that has a colony of convict cichlids (Archocentrus Nigrofasciatus). In March, a large clutch of babies was spawned, and we've been actively working to raise them. The tank is well maintained, it has a lot of filtration, and it has a large capacity air pump feeding air stones (read: a lot of bubbles in the tank background).

Recently the kids were playing upstairs, and as they are oft to do, roughhousing and jumping around. Unbeknownst to them, they knocked the air line from the air pump to the air stones. Earlier in the week, they helped me clean out the tank and do a water change, too. They didn't realize it, but these events conspired to make a fatal flaw in our system apparent.

I came home from work Thursday night and the kids said "hey dad, you should go up and feed the fish, they look really hungry!" I thought this was an odd statement, so I went up and inspected the tank. What I saw was all of the fish hovering at the surface of the water, some on their sides, their gills pumping rapidly. As I looked and saw the tank, I heard the air pump running louder than normal. I looked down and saw that the air hose was no longer attached. I quickly reattached it, which allowed the bubbles (and air) to flow back into the tank. However, this was only part of the problem. I also noticed a tell tale sign of what also wasn't normal; there was no splashing water from the filter return. The water was just running under the surface because the hose had slipped down below the water line. Add to that the fact that we also have about 70 juvenile fish, each now averaging an inch in length or a little more, along with the older denizens in the tank. Individually, any of these would not be a problem. Taken together, though, what resulted was a massive failure to the system.

We had to act quickly, so I hooked up the water siphon to the bathroom sink and drained out about 20 gallons of the tank water. this served two purposes. first, it allowed for agitation of the surface by the filter return, thus creating an oxygen exchange again. This mixed with the bubbles from the air filter would allow the oxygen to return to the water. Second, I turned on the hose and briskly sprayed water all around the tank (the motion of which would help oxygenate the fishes gills and help them to breathe more naturally. Finally, I surveyed the damage after I saw the fish were moving on their own accord and had perked up considerably. The result was 15 dead juveniles. They had suffocated from lack of oxygen.

When all was calmed down, I gathered my kids together and taught them a bit more about the aquarium ecosystem and how it worked, about oxygen exchange, about water column movement and also overcrowding and overpopulation for a small environment. They had asked me why no fish had died when we had power outages. Then not only did the air pump stop working but the water pump and heater had, too. Why didn't the fish die then? I explained that the fish population was smaller then, that there weren't so many babies, and those babies hadn't grown to a size to make a major difference to the resources of the tank. Now however, they were large enough to make such an impact.

With that, we had to make a game plan, and part of that plan is to find a new home for about 70% of the fish, as soon as they are big enough "to let go". To let go means I plan to donate them to various fish stores in the county.  The remaining 30% I will let grow as long and as big as they want to get.

Our software ecosystems are similar. Lots of new features may seem innocent by themselves and in isolation, but all it takes is an environmental change to have something catastrophic appear. we may have little to no idea what will cause that to happen, but one day all of our tests pass and everything looks great, and then one seemingly small change takes place, and suddenly everything has gone horribly wrong. I've had that experience with my own tests, and sometimes it's just a little nudge that sends everything over the edge. So what can we do? Just like I needed to educate my kids about oxygen in the water and surface agitation, I also have to be aware of the unique aspects of what makes my software environment "breathe". I may not be able to know every contingency, but knowing the big and important ones will help to stave off disasters.

Saturday, June 18, 2011

First Hand Experience of an Airline Shutdown

It's funny, over the years, I've seen numerous times where situations take place where travelers are stranded, but I always looked at them with the idea of "oh well, I feel bad for those people, but there's little I can do about it (shrug!)" and go on my merry way. Well, karma caught up to me and my family yesterday, when we were supposed to be heading to Southern California and a trip to Disneyland for our family as a celebration of my older daughter's graduating from 6th grade (and just an excuse to get away for a weekend).

When we got to the airport, I noticed that there was a huge line of people at the ticket counters, and that all of the ticket counters screens were showing the United logo... just the United logo. Nothing else. Ominous. Oh well, no worries for us, we already had our boarding passes, so we just skipped through security and went to our gate... and then we saw the true situation. The United computer system was down. Worldwide!

From here, I resorted to my numerous coping mechanisms. Gather the kids together, pull out the games and the books, pull out my laptop and start editing audio (believe me, when I'm faced with a wait, I always know I can fall back on that). After about four hours of no updates, though, I realized the truth... we weren't going to Santa Ana tonight. No one was flying anywhere on United. What's more, very few people were even able to make alternative plans or book on other airlines. The capacity for flights is maximized; there's little slack in the system. What's more, our electronic methods of surveillance and security do not know how to cope with this. The few flights that left during the evening were those where planes had already landed before the "glitch" and even they had some loud verbal discussion among flight crews as they "paper charted" their courses. Wow, was that an interesting conversation. The flight attendants were genuinely freaked out about a "paper" calculation of a night time flight over the Rocky Mountains. Honestly, I can't say I entirely blame them, but it punctuated the idea that our way of life is indeed in a way held hostage by our reliance on computer systems.

My youngest daughter kept asking me "Dad, can't you help them fix this?!" My daughter of course knows I'm a tester, but I think she has a little too overarching vision of just how powerful her Daddy is (LOL!). I had to explain to her that the problem wasn't here, it was in Chicago. What's more the problem wasn't just affecting us, it was effecting every airport United services, and the ripple effects of this might take a long time to sort out. She kept asking me questions about what could have happened and what could be done to fix it. It felt frustrating to say "I'm sorry, honey, but I really don't know what's going on." For me, that was the biggest aggravation. I can deal with computer shutdowns. I can deal with delayed flights. Heck, I was even prepared to spend the night in the airport armed with my two laptops and a USB stick loaded with books. All of that I can handle. The most aggravating aspect, though was the state of limbo. No one knew what was happening, and other than an intercom message that said "United is experiencing technical difficulties with their computer system. We are working to resolve the issue. We apologize for the inconvenience!", there would be no information.

As a tester, it's the information that I provide that adds value and allows us to make decisions. My problem was I was stuck in a fog. I couldn't make a decision. What was going to happen? Was my flight cancelled? Would it resume later? Do we need to pull the plug and go home (an option we had because we were still on the home leg of the trip. I can only imagine how fun this must have been for people trying to get back home)? It's not the problem, it's the lack of information following the problem that's the real kicker.

I often tell my kids that having a flexible attitude can help in a lot of negative situations. Being rigid in approach and thinking can cause a lot of stress and frustration. It's common to see people lose it when they are delayed. They don't really have a coping mechanism. They just fume, and I saw some people doing that. I liked Steven Covey's idea from "7 Habits" where he said to "always have a Plan B", or "carry the weather with you". The idea here is that I had no control over the flight to Santa Ana. It simply wasn't going to happen. We found that out at 1:00 AM and then proceeded to go home. We have another chance to try again tonight, and with that, we will see if we can work our game plan again. [Update: we managed to all get on the same flight and are now safely in Anaheim at our hotel, across the street from Disneyland. The airport was able to get back to normal within 18 hours of the system getting back online].

We lost half a day, but in that half a day, my older daughter made three new sketches in her art book, my wife made significant progress in a book she was enjoying, my son explored about every inch of SFO, and I completed a full pass of next week's TWiST for first round edits (the beauty of a captive 7 hour wait ;) ). In short, flexibility is needed, and for testers to make the best of the situations they face, they need to know when to zig, when to zag, and when to chuck everything and focus on something else. In truth, I will have to restructure some of my plans. I will now be taking a day off from work that I hadn't planned for, but I'm working today to help cover for that. I lost a day and delayed a vacation for a day. On the flip side, I can only imagine how many millions of dollars United lost through this ordeal, and how much lingering bad will might be sitting inside of various passengers minds. Will they book United again in the near future? Who knows?! Still, it's interesting to see the reactions of people and how they cope with these situations, and realizing that, even when things really don't go your way, it need not be the end of the world. A little flexibility can go a long way.

Monday, February 7, 2011

BOOK CLUB: How We Test Software at Microsoft (14/16)

This is the sixth and final part of Section 3 in “How We Test Software at Microsoft ”. This chapter focuses on how Microsoft tests the large (sometimes extremely large) set of features that they call Software+Services, and the special challenges that come from testing applications that run at mega-scale. this chapter is the physically largest in the entire book and as such, this is the longest of the chapter reviews yet. Note, as in previous chapter reviews, Red Text means that the section in question is verbatim (or almost verbatim) as to what is printed in the actual book.





Chapter 14: Testing Software Plus Services

Ken Johnston wrote this chapter and he opens it up with an homage to a great book, the dangerous Book for Boys.Assembling the ultimate adventure backpacks and finding out what gear is essential and necessary for getting into danger, is a big part of the fun. Some of the stuff was great fun, but some of the areas were a little big hard to get his head around (tripwires are indeed dangerous and hazardous, which of course made them even more enjoyable and fun for his son. the point being, boys and danger go hand in hand, and the The Dangerous Book for Boys helps differentiate between big dangers and little dangers.

Software Plus Services (S+S) is an area that Microsoft is now focusing on ,and just like that Dangerous book for Boys, they also realize that S+S is an area that has its own dangers, both big and small. Services-oriented architecture (SOA) and software as a service (SaaS) is a big topic today, as is anything having to do with software that exists "in the cloud". Currently, though, there is no book that covers the dangers that lurk out there for those looking to take on development of Software+services, though this chapter certainly tries.

Two Parts: About Services and Test Techniques

To make the most of the chapter and help us get our heads around the issues, Ken has given us two sections to look at and consider. Part 1 deals with Microsoft’s services history and strategy, and how it compares and contrasts with regular software applications and Software as a Service (SaaS). Part 2 discusses the testing strategy needed to test services.

Part 1: About Services


The Microsoft Services Strategy

When Microsoft refers to Software Plus Services, what exactly are they talking about? The idea is that distributed software allows for companies to provide various online services (think Hotmail, Facebook, or Google Docs as examples) that leverage the local processing power of the 800 million plus computers out on the Internet (as of the date the book was published; there's likely more than a billion of them now if you include all the smart phone devices).

Shifting to Internet Services as the Focus

In 1995, Bill Gates released a memo that called for Microsoft to embrace the Internet as the top priority for the company. Inside Microsoft, this is often referred to as the “Internet memo,” an excerpt of which follows. With clear marching orders, the engineers of Microsoft turned their eyes toward competing full force with Netscape, America Online, and several other companies ahead of us on the Internet bandwagon.

The Internet tidal wave

I have gone through several stages of increasing my views of its importance. Now I assign the Internet the highest level of importance. In this memo I want to make clear that our focus on the Internet is crucial to every part of our business. The Internet is the most important single development to come along since the IBM PC was introduced in 1981.

—Bill Gates, May 26, 1995

Having been part of the middle wave of the development of the Internet (I can't claim the seventies or eighties, but I can claim anything after 1991 :) ), I remember well how it felt to try to connect a PC to a network and how much extra stuff had to be done to get it on, communicate, and then all of the extra tools necessary to facilitate communication with other machines. Over time Microsoft really took seriously the desire and the focus of providing Internet capabilities; windows 95 made it very easy to network computers without having to go through special steps; plug and play Internet was pretty much available to anyone that wanted to use it. But now that people were online, what would they do with that access?

Older models of "services" already existed. FTP allowed the user to transfer files; UUCP and USENET allowed for communications across various newsgroups on thousands of topics. With the coming of the Web, many of those older structures grew into new sites and new services; while USENET still exists, and the groups and hierarchies are still out there, but when we talk about services today, again, it's the Facebook's, the twitter's and the Tumblr's out there that people first think of.


In October 2005, another memo was sent out to the Microsoft staff. The “Services memo,” In it, a word of seamless integration was discussed, where applications and online services would not just be complementary, but nearly indistinguishable from each other. This is where Microsoft's strategy of SoftwarePlus Services comes from, along with their goal of integrating the Windows operating system and the various online offerings (including SaaS and various Web 2.0 offering).


The Internet Services Disruption

Today there are three key tenets that are driving fundamental shifts in the landscape— all of which are related in some way to services. It’s key to embrace these tenets within the context of our products and services.

1. The power of the advertising-supported economic model.

2. The effectiveness of a new delivery and adoption model.

3. The demand for compelling, integrated user experiences that “just work.”

—Ray Ozzie, October 28, 2005

S+S goes farther than simple having a PC on the network. PCs and mobile devices are all part of this equation, and the list of supported products is growing. Microsoft has also embraced "cloud services" and to that effect have released a software platform called Azure (www.microsoft.com/azure). Azure will be a focus for such technologies as virtual machines and "in the cloud" storage (think of Dropbox or again, Google Docs as examples). Live Mesh is a part of Azure that allows users the ability to synchronize data across multiple systems (computers, smart phones, etc.). Smug Mug and Twitter are active users of a cloud bases system that amazon offers called Amazon Simple Storage Service (S3). On the flip side of this, if S3 goes down, the services offered go down, too (hey, nobody rides for free :) ).

Growing from Large Scale to Mega Scale

Microsoft launched the Microsoft Network (MSN) in 1994, and it quickly became the second largest dial-up service in the U.S. MSN was a large service, requiring thousands of production servers to operate.

With the acquisition of Hotmail and WebTV, Microsoft went from a big player in the online world to a truly massive player. WebTV brought to Microsoft the idea and the need for "service groups", which is where large blocks of machines work mostly independently of other blocks. Service groups have the added benefit of not bringing down an entire site or service. If one group goes down, other groups will still be running to keep the service going (if perhaps not at the same level of performance as with all service groups up and active). Hotmail taught Microsoft of the benefit of what have come to be called "Field Replaceable Units" (FRUs) from Hotmail. These FRU systems were little more than motherboards and a hard drive and a power source, and they were on flat trays like open pizza boxes. As is pointed out in the chapter, this was in the days prior to the wide scale growth of the gigahertz plus multicore CPU's that required skyscraper heat sinks to keep cool, and just using the server rooms A/C would be sufficient to keep the systems running optimally (ah, those halcyon days ;) ).


These concepts are still relevant and in practice, but there are limits; at the time of the writing of this book, Microsoft was adding 10,000 computers each month to its datacenters just to meet demand. Even with a modular open pizza box design and structure, that's a lot of hardware that requires electricity and cooling (which, of course, likewise requires electricity) as well as other space and infrastructure needs (buildings, rooms, cabling, security, etc.). A few years ago, Microsoft standardizaed on data center racks loaded with all the gear they would need, and these racks were dense. Now, instead of just loading racks, Microsoft now develops Datacenter Modules that are the size of freight shipping containers, and referred to as container SKU's. For visual purposes, think of the size of an 18 Wheeler's container hold. Now picture that thing loaded with racks, and with large connector harnesses. Now picture each of those being dropped somewhere and connected to hook up and add to a datacenter. That gives you an idea of the scale that these S+S systems are using and adding routinely, as well as the size of the systems they rotate out when maintenance or retiring of machines is required.



Microsoft Services Factoids

Number of servers: On average, Microsoft adds 10,000 servers to its infrastructure every month.

Datacenters: On average, the new datacenters Microsoft is building to support Software Plus

Services: cost about $500 million (USD) and are the size of five football fields.

Windows Live ID: WLID (formerly Microsoft Passport) processes more than 1 billion authentications per day.

Performance: Microsoft services infrastructure receives more than 1 trillion rows of performance data every day through System Center (80,000 performance counters collected and 1 million events collected).

Number of services: Microsoft has more than 200 named services and will soon have more than 300 named services. Even this is not an accurate count of services because some, such as Office Online, include distinct services such as Clip Art, Templates, and the Thesaurus feature.


Power Is the Bottleneck to Growth

Moore's Law is still proving to be in effect. Roughly every 18 months, the processor speed, capabilities, storage and RAM availability just about double. Along with that, the need for power, cooling, and sufficient space likewise rises along with it. The average datacenter costs about $500 million to construct. To this end, Microsoft is working aggressively with vendors and Other Original Equipment Manufacturers (OEM's) to find ways to help design systems and datacenter infrastructure that is the most efficient with the way that it can draw power and be constructed to get the most bang for the buck in as many categories as possible.
In a sidebar presentation, Eric Hautala makes the case that "Producing higher efficiency light bulbs is a fine way to reduce power consumption, but learning to see in the dark is much better..." The point of the comments is that we can work towards making systems more efficient, or we can work towards making the code and applications themselves more efficient”. In other words, rather than try to accommodate for a greater and greater need and find ways of doing the same job cheaper, how about finding ways to do this and reduce a particular need entirely?



Services vs. Packaged Product

When we refer to a product that is purchased on a CD or DVD, that's a "shrink-wrap" product, and typically comes with all of the items associated with that kind of delivery, including a box, a cellophane wrapper, and glossy artwork to make the package appealing. Same as when the applications are installed on a computer that is being purchased. The distinction is blurring, though, as many examples of items are being purchased as Shrink-Wrap software (Xbox games, Office versions, etc.) but they also access content and service items online that can be downloaded and used to enhance the product, or to connect with the Internet to update and interact (Xbox Live, for instance). the rebranding of Hotmail and passport to be part of the Windows Live brand are movements towards a more defined services model on the web.

Windows Live Mail is currently the largest e-mail service in the world with hundreds of millions of users. WLM works with many different Web browsers, Office and Outlook versions, PC's and Smart devices. Many of these services require a live connection to the Internet (example, you can't get new email when you are not connected to the Internet). By contrast, Web 2.0 tends to rely on newer browsers, Flash, or Microsoft SilverLight to allow for processing changes to be moved away from the web server to applications and modules that are built into or rendered and processed in the browser, and subsequently, through the client device itself. Because of this ability to have the local system do the processing, in many cases, these systems can be offline to do this.

Moving from Stand-Alone to Layered Services

Generally when a web site was developed in the early days of the net, sites were self-contained, and most of the processing or validation or workflows were programmed on the server and the users relied on that server being available. No server, no site, no service.

Layered services allows the ability for companies to leverage multiple machines, including the client machine itself, to help process transactions, track progress, and select different workflows and execute those workflows. In truth, most sites that look like they are single site applications are actually layered and spread out systems.


Using eBay as an example, eBay acquired PayPal and then integrated it with the eBay experience. Not only is it a service for eBay, but for anyone who wants to request and send money. Standalone services generally are easier to test by comparison to those that are leveraged and spread out through many different technologies. When a system is layered, the testing requirement can go up significantly (even exponentially).

Part 2: Testing Software Plus Services

In this section, Ken goes through describing test techniques that can be used for testing services, and also walking through various approaches to using those test techniques.

Waves of Innovation

Microsoft has gone through a number of innovations in the way that computers, and the way people interact with them has changed over its existence. The evolution of these methods follows below:

1. Desktop computing and networked resources
2. Client/server
3. Enterprise computing
4. Software as a Service (SaaS) or Web 1.0 development
5. Software Plus Services (S+S) or Web 2.0 development

With each change, software testing has had to evolve and change along with it so as to meet the new challenges. Number of users and methods of interacting have blossomed and in some cases exploded. With each wave, most of the older tests continue being used, and more and newer tests get added.

Designing the Right S+S and Services Test Approach


Ken brings us back to the Dangerous Book For Boys idea, and how he wishes he could write the Dangerous Book for Software Plus Services. In his mind, this book would have a lock on it but no key. Readers would have to figure out how to pick the lock if they want to read the actual secrets within the book. Ken confesses his reason for this is that, were he to write the book, he'd have to share all the mistakes he has made developing, testing, and shipping these services.

Client Support

It's not just browsers to deal with these days. For many services, testers at Microsoft have to look at a matrix of browsers, plus various applications (Outlook, Outlook Express, and Windows Live Mail client for the example of Mail services)as well as mobile devices. Add to that other products that interact with Mail services and the various languages and regions that have to be considered. The point is that exhaustively testing these options is impossible, so the testers have to determine various aspects to help them pare down the matrix. Market share is considered for various browsers and clients, as are the goals of the particular service and which area they are trying to serve. A risk analysis needs to be made to help determine which areas will require a large scale push for testing and which areas can be given a lower priority or even pushed down or of the list completely.

Built on a Server

The benefit of having many of the services integrated into a Server platform is that many of the systems have already been tested and integrated into the systems. Many of the services and their interactions with the underlying server code have been tested over several years as the applications have been developed. As the service matures, many of the issues are located within areas like performance and enterprise wide actions (such as having multiple people accessing the same objects at the same time, and how they interact under those circumstances).

Server products also require testing and a focus on how the product can be managed and how the service can actually scale and just how far the system can be scaled. The simpler and more hands off the management of the systems can be, the more profitable and effective the services can be, as they will require less overall direct interaction and more automated or scripted interaction will help the environments scale at a lower cost point.


The testing efforts are focused on how the users and remote applications access the service via API or direct interaction. Integration testing is also critical and how many items can be tested (since not every possible combination can be tested; the order of magnitude is absolutely massive).

Loosely Coupled vs. Tightly Coupled Services

Layered services are adding new features and growing with new dependencies with every release. Coupling is another way of saying how dependent one program is upon another. When systems change frequently, loosely couple systems are desired. With loosely coupled systems, the fewer the dependencies, the easier it is to innovate individual areas and add features for specific areas. More tightly coupled services require closer monitoring and more project management. To ensure that features are delivered with high quality and to ensure that a change in one component doesn't cause ill effects in other areas. Loosely coupled services are important when integrating with third party applications or external devices and services. When dealing with credit card companies or payment services, loosely coupled systems were easier to implement and, subsequently, ship. More tightly coupled services were much less easy to implement and often prone to delays in shipping.

Stateless to Stateful

Stateless services, such as sending an email, or loading a series of web pages, don't require that data from one transaction be stored and compared for other transactions. Stateful transactions, however, depend on the data from each transaction to carry over (logging into a secure service, making a credit card payment, etc.). When stateful events are disrupted or cannot be completed, there can be significant impact on the user experience. When a service takes a long time to complete a transaction and needs to store unique user or business-critical data, it is considered to be more stateful and thus less resilient to failure. Compare stateful to working on a Microsoft Word document for hours, and then experiencing a crash just as you are trying to save the file. It might be a single crash, but the impact on the user is dramatic.

Time to Market or Features and Quality

There's a fine balancing act between being first to market and having a product that can maintain the hold once a product gets there. While often a product that is first in the market can have a dominant position for a long time, there's no guarantee that will happen. The example of Friendster is used as one of the first Social networking sites, but MySpace and Facebook were able to jump ahead because of features and quality (and then mySpace was eclipsed by Facebook as the dominant social media app, and a few years down the road, provided Facebook doesn't keep innovating and working to make the platform more appealing, perhaps another innovator will take its place). When balancing between the two, both need to be considered, but if the time to market is all that's being considered, it's a good bet that someone else will take over your spot because they offered better quality and value that trumped your first arrival innovation.

Release Frequency and Naming

Hand in hand with the time to market and quality question is whether or not to release regularly with regular updates, or to offer the software as a beta version publicly. Google did exactly this with Gmail, allowing it to incubate and get feedback over the course of three years. Releasing it as a public beta alerted users that there very likely would be problems with it, and in doing so, they were able to solicit feedback and develop some mind share without fearing that the bugs that were found would be such a turn off that many would stop using the service. I know that this is often a plus for me, when I use software for various purposes. I am OK with some bugs in these types of environments, provided they don't interfere or cause problems with "mission critical" applications that I may be working with. Released software that's been "polished" (or purported to be) I'm a little less forgiving with, especially if the bugs are frequent or very up front (odd esoteric things I can deal with, especially if there are identified workarounds).

Testing Techniques for S+S

So now that we have identified the potential "Danger areas", what can we set up to actually test them?

Fully Automated Deployments

Ideally, in a Microsoft environment and when working with Microsoft products, the testers work through fully automated deployments, where the interaction of the tester or developer with the installation and configuring of the software should be minimal. The more hand holding required, the farther away from a final completed solution the product actually is. Here are some tips offered by Ken and his team:


Tip: If operations has to do anything more than double-click an icon and wait for the green light to come back saying the deployment completed successfully, there are still bugs in the deployment code. They might be design bugs, but they are still bugs.
 
Tip: If the deployment guide to operations has any instructions other than, “Go here and double-click this file,” there are bugs in the deployment guide.
 
Tip: If there needs to be more than two engineers in the room to ensure that the deployment is successful, there are design flaws in the deployment code and the deployment guide.

Great deployment is critical to operational excellence

What makes for a good deployment scenario? According to Ken, the key features of a good deployment are:

- zero downtime

- zero data loss

- partial production upgrades (a service might upgrade a small percentage of what is considered the production servers to the new version of the code)

- rolling upgrades (a service can have portions of the live production servers upgraded automatically without any user-experienced downtime)

- fast rollback (the safety net used if anything goes wrong; if something goes wrong, roll back to a previously known good state)

Test Environments

The simple fact is that, with Microsoft products and Microsoft services in particular, there is no one size fits all approach. There can't be. Each service and each setup may require more or less interaction, more or less resources, and demand will ultimately change and determine how a service acts and is interacted with. In many cases, different test environments must be set up to test as many parameters and features as possible.

The One Box

When you have a single test machine, with all of the required software running on one platform, and all of the code needed to run the service and system exists on one physical computer (or VM) this is a one box setup for testing. One box has the advantage of allowing fairly quick read/write transactions and interactions between components. The systems can all be checked from the same terminal service session, and configuration changes to one have the effect of changing performance to another option. RAM and disk access are all specific to that machine, and if one piece suffers, everything suffers. Still for speed and debugging purposes, having a one box test setup can be and often is one of the single most important and quick to debug and troubleshoot test systems there is.

The Test Cluster

A Test cluster, by contrast, may be several physical or virtual machines running simultaneously and interacting with one another. While these are not as easy to diagnose and debug issues as the "one box" systems are, they definitely do provide for a more realistic environment, one where multiple security groups can be sent and where context can be tested and privileges can be across a network or specific to a particular machine. An example is when a datacenter or an enterprise customer has a web server, a database, server, a file server and a terminal services server. In a one box environment, these are all part of the same system. By contrast, in a real production environment, to prevent from a single point of failure, the web server would be on one system, the file server on another, the database system on another, and the secure remote access on yet another machine (or at least being managed by another machine). These options allow the user to see the context in which successful connections are being made, and in those where they are not so obvious).

The Perf and Scale Cluster

When we use a single box setup, we can do some basic performance analysis of the application a d response time with specific criteria easier than we can if we have multiple machines with specific roles. It's just easier and there's no network latency or other aspects of the load of the network really getting in the way. However, some services just place such a demand on a system that the split out systems are required. In addition, it helps to see which components get slower independent of other variables (is the latency due to the web server, the database, or something else?). In many cases, the data load itself can be an issue. Ken uses the example of an email account, and how it will perform differently when there are just ten messages in the inbox vs. thousands of them. By comparison, how does a system work when it is handling email accounts for 10 people or 10,000 people. By determining the performance curve and seeing where the performance impact takes place, it's possible to see how many machines are required and what resources for each so that a "break point" is determined, where one physical machine needs to be augmented with two, four, eight and so on machines to meet that service's requirements to scale effectively.


The Integrated Services Test Environment

When there is a collection of services being offered, components may work well in isolation, but there are challenges when all of the services are tested together and tested to see how well they work with each other. One can look at the most recent changes made to windows Live Hotmail to see how much of a challenge this could be. In previous revisions of Live Mail (Hotmail) users were able to open Word Documents, spreadsheets and presentations by downloading the files and opening them via their local copy of Office (that is, if they had it, or they had an equivalent application that could load it. The most recent version of Live Mail now has the ability of opening a stripped down version of various Office applications to view, edit and save files in "real time" through the email account. While I can enjoy the end result, I can only imagine the technical challenges that were required to get a "web specific” version of Microsoft Office to work within Live mail.

The Deployment Test Cluster

Using multiple machines helps the tester take a look at each of the components in turn, and how they would interact with other components were they deployed on different systems. Early in the development and testing process, these tests could be performed on lower end hardware or in a Virtual Machine cluster. As the product gets closer to its release date, more emphasis on creating and running tests on production grade machines would be needed.

Testing Against Production

In any environment, while detailed and specific testing can be made on the most controlled of systems, the real work and real world demands are going to be very different from the automated tests that are performed and even the very specific targeted tests that are run on a number of systems. For the true impact of changes and new development to be seen, there must be a level of testing performed that covers the real system. Yes, that means the live production environment. While that doesn’t mean that the service must be brought down entirely to deploy and test, it does mean that many of the components will need to be loaded and actually tested in real world scenarios on the live systems. Examples of this can be seen with Facebook; they actually take a feature and roll it out to a small subset of their customers to see how it responds to real interaction and real usage. Based on those tests, they can choose to roll it out further or bring the component back in house for continued evaluation, testing and possible reworking.


Production Dogfood, Now with More Big Iron
 

The idea of a “dogfood” network or system was discussed in chapter 11, and it has to do with the idea that a company will live with its service before it is “inflicted” on anyone else. The developers and testers and the other folks who work at Microsoft are the customers before their customers are. Having worked at Cisco Systems in the 1990’s, I am familiar with living on a dogfood network and having to often check and see what is happening; it’s a sobering situation, and one that helps users realize all of the things they need to be aware of that might impact their customers, as well as bring to light needed changes that could help the product and the user experience overall.

Production Data, so Tempting but Risky Too
 

Microsoft has a collection of tens of thousands of documents that have been cleansed (sanitized) so that the information inside the documents cannot get out and be seen as a security risk, yet can still be used as a good representation of real production data. Having access to this may documents can again help the developers and testers see how real data in real situations can impact performance for applications and services. While having a set of 100,000 email addresses that are live could potentially create a problem if a script is mis-configured (lots of unintended SPAM email getting sent), having email addresses that are nonsense but still form a valid address construction can be helpful in analyzing how a database stores them and how long it takes for those addresses to be served up and used.

Performance Test Metrics for Services
 

Many of the metrics used for testing performance of applications can also be used to test services. Some tests are specific to services and are most helpful in that particular paradigm:

- Page Load Time 1 (PLT1):
Measures the amount of time it takes for a browser to load a new page from the first request to the final bit of data. A new page is any Web page the browser has never been to before and for which it has no content cached.

- Page Load Time 2 (PLT2): Measures the amount of time it takes for a page to load on every visit after the first. This should always be faster than PLT1 because the browser should have some content cached.
 
- Page Weight: The size of a Web page in bytes. A page consisting of more bytes of data will typically load more slowly than a lighter weight page will.
 
- Compressibility: Measures the compression potential of files and images.
 
- Expiration Date Set Test: Validates that relatively static content has an expiration date greater than today.
 
- Round Trip Analysis: Evaluates number of round trips for any request and identifies ways to reduce them.

Some tools that help to make these test possible:

- Visual Round Trip Analyzer (VRTA)

- Fiddler (http://www.fiddler2.com)


Several Other Critical Thoughts on S+S
 

There are a number of additional areas that are helpful when it comes to testing services. This review is way longer than any other chapter review I have done, so to not draw this out, I have made this section smaller than it probably deserves. Check out the book for more detailed explanations of each of these.

Continuous Quality Improvement Program
 

Just like how Microsoft Office and SQL Server have the option for users to participate in Customer Improvement Programs, so do the service offerings. Plain and simple, the real world usage patterns and actual interaction with the systems is going to be multi-faceted and vary from user to user and company to company. While it is possible to make some assumptions and get some likely scenarios, the fact is that there will be many areas that were not considered or come to light only after being seen in the field. Freqwuent monitoring of the actual usage statistics and pain points that customers actually experience will help to inform future development and testing.

Common Bugs I’ve Seen Missed
 

The simple fact is that, no matter how detailed a test team is, no matter how focused they are and how much of a real world they simulate, the test environment is, at best, a crude approximation of what happens in the real world. This is super clear when it comes to working with Software+Services, because let’s face it, there really isn’t any way to 100% guarantee that all scenarios are covered for a system that has thousands of users, let along hundreds of millions. There will be cases missed, there will be situations never considered, and yes, there will be bugs that come to light only when they have actually been exposed to real world use. Severe bugs and security issues will continue to come to light, and often it takes a concerted effort to learn from the issues discovered. It’s easy to criticize a company for a product that has a “glaring” defect, but again, glaring to who? Under what circumstances? Could it have been foreseen? Could it have been prevented? It’s easy to Monday morning quarterback a lot of testing decisions and look at bugs in hindsight. It’s not so easy when the systems are under development and their full scope and impact may not be seen for weeks, months or in some cases, years, until they are scaled up to the tens or even hundreds of millions of users.

Services are now a huge part of the Internet landscape. For many Internet savvy users, it’s all they access. Think of going online to check your Gmail account, listen to music through iTunes or Rhapsody, communicate with friends via Facebook, watch videos through YouTube or Hulu, or for that matter, track your television shows with SideReel (hey, why not put in a shameless plug ;) ). The fact is, most of the Internet properties out there that people access are related to or are in and of themselves Software+Services offerings. These systems are going to continue to grow in the future, and they are going to become every bit as powerful and every bit as sophisticated as desktop applications that we use and take for granted. Likewise, the expectations that they perform the same way will continue to rise, and thus the need to diligently test them will rise along with those expectations.

Monday, September 13, 2010

When Catastrophic Failure Hits Home


I've wanted, and circumstances have kind of forced me, to take a break from the blog and contemplate on something that happened in my community over this past week. As many of you will see when you look at my profile, I live on the San Francisco Peninsula. More specifically, I live in the small town of San Bruno, CA. San Bruno became a headline news story last Friday and over the weekend all over the country. This was due to a natural gas line that exploded in a residential neighborhood. The explosion caused the confirmed fatalities for four people, destroyed several dozen homes, displaced hundreds of families, and required the concerted efforts of many cities fire crews and emergency personnel to maintain order and prevent the fire from spreading into other neighborhoods.

At 6:15 PM on Thursday evening, September 9, 2010, reports came in of a loud explosion in San Bruno near Crestmoor Canyon. Due to the proximity of San Bruno to the San Francisco International Airport, many thought that a jet airliner crashed in the canyon, and that the fire we were seeing was from such a crash. Shortly afterwards, we heard a report that all planes were accounted for, and that the odds of it being an airplane were non-existant. Then what? What was causing a massive fireball to shoot up in the air to heights of 100 feet and higher?

We didn't get to have the time to speculate on exactly what was happening, because our first motivation was to pack up our things and get out of the area as quickly as possible. Actually, this was more my wife and kids decision, because I had not returned home yet. I first learned of this situation when I got off of the Bay Area Rapid Transit (BART) train that I take every day to come home from work. As I drove up the street heading for my home, I saw the smoke and I saw the fireball, and it looked from my vantage point to be in the canyon. This scared me in a big way; Crestmoor Canyon is filled with Eucalyptus trees. Eucalyptus trees were what fueled the Berkeley Hills fire of 1991, in which 3000 homes and other buildings were destroyed. If Crestmoor Canyon were indeed on fire, and those Eucalyptus trees were to catch, then a large part of western San Bruno had the potential to go up in flames.

As I fought my way home through traffic and a police barricade and being told I could not drive into my neighborhood because an evacuation was in effect, I parked my car on the onramp to Southbound 280 and made my way by foot up to my wife's parent's place, where she and the kids were staying. From there, I was able to get the rest of the story. The fire was not in the canyon, but was in the neighborhood just to the west of the canyon, often referred to as "Crestmoor 2" by real estate agents, or Glen View/Claremont by two of the main streets. The cause was determined to be a 30 inch gas line that had ruptured, and it was the high pressure in this gas line that made the explosion so severe that it created a crater 40 feet across and 40 feet deep, while registering a 1.3 quake on the Richter scale as reported by the US Geological Service. The explosion took place very near the intersection of Glen View and Claremont, which is a low lying section of road locals call "the gulley" and where storm and water runoff heads into Crestmoor Canyon. The fireball that resulted from this explosion quickly incinerated the houses immediately surrounding it, and then the fire spread to other houses. In addition, the blast also damaged a water main, which meant there was no immediate water they could use to fight the fire; they had to either fill trucks some distance away or run hoses from hydrants more than 1/4 mile away.

Firefighters from several nearby cities were called in to help fight the blaze, work with the city and county to shut down the gas line that was feeding the fireball (and the fact that it had to be shut down slowly so as to not cause ruptures further down the pipe) plus a committed team of firefighters that were told, in no uncertain terms, "do whatever you have to do to make sure that fire does not hit the canyon!" They knew that, if the fire did spread to the canyon, many other neighborhoods would be in danger.

My family spent the night on the floor huddled around a battery powered radio (since the power was out for the greater western part of town), listening to the reports. With the smoke and the confusion, we were hearing estimates of 50 + houses destroyed and 120+ houses damaged. Knowing how many of our friends lived in that neighborhood, we knew this was not a good sign. With the power being out, we also had a hard time communicating with our friends and family to let them know we were safe. What I later found to be cool was the fact that a friend we did contact made a point to post to both my and my wife's Facebook pages that we were safe, had been evacuated, and that we were dealing with no power and an overloaded cell phone network and otherwise couldn't call or email to let people know we were OK. We later commented that this was the first "Facebook" enabled disaster, and that Facebook played a major role in helping get the word out to people about the event (expect me to do a post about that another time).

As the evening wore on, I decided that I needed to get my car off the freeway, and so I drove and found a place to park. As I was coming back up to where my wife's parents lived, I was stopped by the police and told I couldn't continue. I pulled out my wallet and said "look, I live here; my wife and children are just up the street. You may escort me up if you wish, you may arrest me if you wish, but one way or another, I'm getting back to my wife and kids!" With that, they nodded and let me pass and get back to my in-law's home.

Friday morning, we received the news that the evacuation order for our neighborhood had been removed, and that we could go back to our homes. We were overjoyed to find out that the fire had been contained to the Claremont area, that the fire did not make it down into the canyon and that our house was unaffected. With this brief joy, however, came the stark realization that many of our neighbors had lost everything. With power restored, we could see on the television and the Internet the pictures and video of the damage. Several dozen houses were gone. Not just burned down, but vaporized! Nothing left but chimneys and foundations. To this we received word that several people lost their lives, one of the fatalities being the older sister of a boy in the Cub Scout Pack that I used to lead, and that hundreds of families were displaced, for how long was anyone's guess.

As a software tester, I can't help but look at this in the lens of what I do for a living. I heard reports from people that said that due to a "glitch", the pipe ruptured. When a pipe ruptures and close to 50 houses are destroyed, and four people are confirmed dead, that is not a "glitch", that is a catastrophic failure! Many things have come to light from this incident that I had not known about my neighborhood, such as the fact that the Claremont area was sitting atop this 30 inch gas line running right through a densely populated neighborhood. I do not think any of the residents knew that. I certainly didn't, and I've lived here for 11 years. I did do some research to find out where the rest of these large scale pipes were located, and thankfully, the one that services my neighborhood is out by the main road that leads up to our neighborhood, but no houses are near it. Questions will certainly be asked, most of them pointed at Pacific Gas and Electric (P.G. & E.). During the reports, there were comments that residents' smelled gas for three weeks prior to the explosion, but little was known if anything was done about it. If indeed P.G. & E. did know there was an issue and did not do anything about it, then this rests squarely with them. Bigger questions will rise, such as "was the pipe decayed because of its location?" or "did running groundwater weaken the integrity of the pipe?", but I think one of the questions needs to be "who decided that building a densely populated community right over such a large gas main was a good idea?". Perhaps this is common, and I just wasn't aware of it before. Knowing what I know how, had I been in the market for a house, I'd think twice about buying a house where a primary 30 inch gas main lies just under my development.

Again, as a tester, I often think about these things in the abstract; a software failure that is deemed catastrophic is one where the system cannot perform its function. Even in such circumstances, the outcome rarely consists of more than lost time or loss of revenue (important in their own spheres, to be sure). In this case, a catastrophic failure has resulted in the deaths of four people (perhaps more), the destruction of dozens of homes and the displacement of hundreds of families. One thing is for certain, it's unlikely I'll be forgetting about this catastrophic failure any time soon!