Thursday, December 11, 2014

The Automation Pyramid

The Automation Pyramid

Published: October 10, 2013 by Brian Kitchener

Across the world, in many automation models, the concept of the ‘automation pyramid’ is spoken about. The idea is fairly simple. There are multiple layers of automation, each focusing on different areas of the application, and offering a different degree of coverage. The base of the pyramid are the Unit Tests, these are executed against the code. Next is the API test, executing against a service layer. Finally, at the top of the pyramid sits the UI tests that actually validate the application as a whole. Each of these layers offers a different level of coverage. Unit Tests are fast and can test many different permutations, but don’t test any integrations between components. API tests can validate our units integrate, but can’t test an end-to-end user scenario. UI tests validate multiple components in a single test, but are typically very fragile and take a long time to run. So each layer provides a pivotal role, and needs to be tested thoroughly.



The Base of the Pyramid

At the base of the pyramid are the Unit Tests. As our foundation, they will have the broadest coverage, and should test as many permutations as possible. Typically they will test a specific “Unit” of code. They will instantiate a class, call some functions, and verify the functions return the correct results. Unit Tests are built in the same language as the code and will usually be stored and executed with it. Because they are executing a piece of code, it’s extremely easy to test for many different permutations. For example, if I have a function that does some simple arithmetic, I could test it 100 different ways in a couple seconds. It could take hours to verify the logic using the UI, because of all the navigation and setup required. Because of this, our unit tests should provide extremely broad coverage, testing as many permutations as possible. They should attempt to test and validate important business logic at a component level. However remember that they are unit tests, and will not test how components work together, only that they individually work as expected.

The Middle of the Pyramid

The next level of our pyramid is made up of API/Service Layer Tests. These should fit somewhere in between UI and Unit tests in complexity and scope. This layer is where we first start to test how our components integrate together, and how they handle real data. The advantage of API testing is that a lot of logic can be validated without being dependent upon the UI. For example, let’s suppose we are building a new web site. Even before the UI is finished, the service layer and authentication mechanisms could be finished. I can build a set of API tests that verify the login service works as expected. And while the tests aren’t quite as fast as Unit tests, you can expect most API’s to return a result in less than a second. That means that testing for a variety of conditions such as invalid credentials, timeouts, or special characters is very easy and very fast. This frees up our UI tests to focus on testing the UI and end-user functionality, instead of trying to validate every piece of business logic. Without our API tests, we wouldn’t even be able to start testing until the UI is finished.

The Top of the Pyramid

The highest level, and therefore the smallest layer, is comprised of the UI tests. It is at this level that we actually launch our application and perform an end-to-end test through it. It is at this point when we finally verify that all the individual pieces work together as a cohesive whole. But since much of the business logic has been covered in the lower levels, the tests can focus on validating the UI looks and behaves appropriately. It’s important to remember that UI tests are extremely slow in comparison with our UI and API level tests. A UI test will take usually a minute or so, when a API test takes several seconds and a unit test takes less than a second. But a single UI test will validate hundreds of components in a single test, so its coverage is much deeper than the other layers. However the UI is constantly changing and being updated by new functionality, which means that the UI tests are inherently less stable, and will require more maintenance than the other layers. This is why we try to limit our testing in this layer to as little as possible.
failed pyramid
By breaking the application apart into these layers we are able to systematically and logically break apart the application and test it in components. Without a strong foundation of Unit and API tests, we are forced to automate every possible scenario in the UI layer. And while it’s still possible to automate only UI tests, our resulting automation suite ends up being expensive, slow, and fairly fragile. And while it’s possible to balance a pyramid on its end, it certainly won’t be stable. It is for this reason that many companies fail at efforts to implement effective automation strategies. If we take a systematic approach, and all the levels are automated to an appropriate degree, we can build and maintain a cohesive suite of automated tests with much less effort.

Tuesday, April 8, 2014

Introducing Golem, an Object Oriented C# Framework

Hello everyone, I am pleased to announce the release of Golem, an open source, object oriented C# framework, now available on GitHub.   Golem was an internal tool that ProtoTest has successfully used on a number of our clients' projects, and we like it so much we want to share it with everyone.  It supports a number of automation tools like Selenium-WebDriver, Appium, Microsoft's UIAutomation, and can even test REST services and validate HTTP Traffic.  Tests are written in Visual Studio, using MbUnit, and are executed using Gallio. It's an all-in-one automation tool for anyone working in a .Net environment.  Golem makes building clean, robust, reusable, and scale-able automation easy.  It's available in NuGet now!

Golem has a number of advantages :   
  1. Simple, Object-Oriented API
  2. Advanced features (data driven testing, parallel test execution)
  3. Robust Reporting and Logging 
  4. Multiple Automation Tools Supported
  5. Fully configurable through an App.Config or through code.  
You can find our official announcement below : (Copied from ProtoTest's Blog)


Prepare yourselves; Golem is coming. A creature from myth and legend returns, reborn for the new age. In ancient folklore, a Golem is a mindless automaton, an unstoppable force, yet it obeys those bold enough to command it. And therein lies the danger, for it will perform any command given to it, faithfully and without rest. For it has no mind of its own, it requires an intelligent being to control it. But beware, the ancient proverb ‘He who rides a tiger is afraid to dismount’ has never been truer: holding power is both addicting and hard to relinquish. Once the power of Golem is yours to command, you cannot go back.
 

Introducing Golem

At ProtoTest, we are passionate about the value of test automation. In fact, our motto is: “Automation makes humans more efficient, not less essential.” We think it is a great way to supplement any manual testing effort. Yet we see a number of people (and companies) struggling to implement test automation in a meaningful way.
There is a qualitative leap between recording and playing back a test and building an enterprise-scale automation suite maintained by dozens of people. This is where a test automation framework steps in. It helps to simplify the process of building, maintaining, and executing a large set of automated tests.
Most automation frameworks have three goals: to simplify the process of building tests, to help diagnose why tests fail, and to allow us to share and reuse code. And while there are a number of automation frameworks on the market, we could not find one that provided the level of simplicity, reusability, and elegance we wanted.
So, we decided to make our own. For the past several years, we have been building and tweaking our own framework in C#. We built the framework around MbUnit and Gallio, because they have advanced features for UI based automation. We included support for a number of tools like Selenium-WebDriver, Microsoft’s UIAutomation, and Appium.
Golem supports most commonly tested enterprise platforms: web browsers, mobile applications, Windows applications, HTTP traffic, and REST services. Tests are written using the industry standard, ‘page object,’ design pattern, and the test report includes as much diagnostic information as we could gather.
In addition, we added several advanced features like data driven testing and parallel test execution. Then we made all of it easily configurable. And now, we want to share it with the world. We are officially announcing the release of the Golem open source project, developed by ProtoTest.
The user group : https://groups.google.com/forum/#!forum/prototest-golem
The source code is available on GitHub: https://github.com/ProtoTest/ProtoTest.Golem
The package is available from nuget: https://www.nuget.org/packages/ProtoTest.Golem/
The documentation is available here : https://github.com/ProtoTest/ProtoTest.Golem/wiki

Tuesday, April 1, 2014

3 Cool things you can do with Javascript Injection

Learning how to interact with or modify an application directly is one of the more advanced things to learn as a tester.  But once we have a hook into an application we can do a variety of things: accessing internal variables, calling methods, manipulating the application’s state, or even modifying the code ourselves.  For some types of applications this is extremely hard, for others it’s relatively easy.  One of the easiest types of applications to manipulate directly is a web page.  This is because most of the code is stored locally on the client, none of it is obfuscated, and all of it is modifiable.  This means that there are a variety of fun things we can do through any browser with a console like Firefox or Chrome. 

You can get to the console in Firefox or chrome by right clicking, selecting inspect element, and clicking on the console tab when the new panel opens.  Any commands we enter into the prompt will be fired against the web page.   Just remember that if a new page is loaded, or if the browser refreshes, anything you do is lost.    Alternately, many automated testing tools provide a way to execute code against the page, and this provides an easy mechanism to manipulate the application automatically. 

1) Modify or execute the code

One of the most useful things to do when testing a web site is to access its internal variables and methods.  This allows us to modify the web page even without a UI.  For instance, suppose that after 60 minutes of inactivity the user is supposed to get a prompt asking them to stay logged in.  We certainly don’t want to have to let our computer idle for an hour every time we test this.   Instead, we can set the web page to timeout after 1 minute by changing the time from 60 minutes to 1.  You can typically ask a developer what the variable is called, or use the console to try to find it.     
So let’s assume a developer told us that the variable was named timeoutMin.  Modifying it is easy.  From the console enter:  document.timeoutMin = 60;  Now the web page should use the new value instead of the old one. 

2) Hiding / Showing Elements

Occasionally when working with a web site that is under development, something won’t display correctly.  For example, an extra panel appears covering the web page.  This prevents you from being able to do any work.  Or perhaps the login panel doesn’t appear.  We are prevented from testing any functionality that requires a login until that issue is fixed. 
Hiding or showing elements is easy, provided they have an Id, class, or name.  Using chrome, you can right click on the element, and select inspect element.  If the element’s HTML contains an id attribute you can use it to manipulate the object.  If it doesn’t have an id, you can even add one, or try getElementByClassName, getElementByName, or getElementByTagName.
The command to hide an element :  document.getElementById(“idOfElement”).style.visibility='hidden’;
The command to show an element :  document.getElementById(“idOfElement”).style.visibility=’visible’;

3) Adding / Removing Page Events

A web page works by registering functions to happen when certain events happen.  There are a variety of different types of events that are called whenever the user clicks, types, or moves the mouse.  For example, a button can have a function registered to the click event called “onclick”.  When the user clicks on the button, the function is called.  If we want we can add, delete, or replace these events with our own.  Let’s look at three examples:
1) To illustrate how to replace an event we will try to disable all click events on the page.  To achieve this we replace the document’s onclick function with one that does nothing. 
document.onclick = function() { return false; };
Most click actions on the page are now disabled.
2) If we don’t want to disable them all let’s add an additional event without removing the old one.  We do this by adding a new listener.  We can add an event to either the entire page, or to a specific element.  For example, let’s suppose I wanted to highlight the element that my mouse is over. I’m going to add two listeners, one to highlight an element under my mouse, and one to un-highlight when the mouse leaves. 
document.addEventListener('mouseover', function(e) { e = e || window.event; e.target.style.border='3px solid red}, false);
document.addEventListener('mouseout', function(e) { e = e || window.event; e.target.style.border=''}, false);
3) Lastly, let’s suppose I wanted to add an alert message when I click the login button.  This will “Pause” the web page and allow me to inspect traffic, html, etc. 
document.getElementById(“idOfElement”). addEventListener(onclick, function(e) { e = e || window.event; }, Alert(“Element was clicked”); false);

 As you can see, there are a variety of reasons why we might need to modify a web page.  It’s not the sort of thing that will needed every day, but is a great extra tool to be added to any SQE’s tool belt.  

Wednesday, February 22, 2012

Using Reflection to track page object actions

Some quick code I thought I would share.  This will use .Net reflection to return the Page Object function currently being executed.  It's just a nice way to track or log what action failed without having to resort to interception or some other technical wizardry.


 public static string GetPageObjectFunctionName()
        {
            string functionName = "";
            System.Diagnostics.StackTrace callStack = new System.Diagnostics.StackTrace();

            System.Diagnostics.StackFrame[] frames = callStack.GetFrames();
            foreach (System.Diagnostics.StackFrame frame in frames)
            {
                if (frame.GetMethod().ToString().Contains("PageObject"))
                {

                    functionName = frame.GetMethod().ReflectedType.Name + "." +  frame.GetMethod().Name.ToString();

                }

            }
            return functionName;
        }

Getting Console Errors From the Browser

So being able to grab Console errors out of the browser has long been a feature I wanted to add into our selenium framework.  Thanks to the help of a couple devs, I've finally worked out a nice way to grab console messages.  Basically the functionality works as follows:


First we override the Console object with some custom functionality every time the page loads.  We injected this into our Page Object base class, that all pages inherit from. This causes it to be called every time the page changes.  After we perform some actions on the page we check our object to see if any console messages have been seen.


There are several ways to do this, but we wanted to be able to differentiate between errors, warnings, logs, and uncaught exceptions.  This code will create the following variables on the page you can grab:

window.console.errorsJson (Errors)
window.console.warnsJson (Warnings)
window.errorsJson (Uncaught Exceptions)
window.console.logsJson (Logs)

First we need to override the console.  This is our magic function that performs all the work on the page.  The beauty of this function is that you can keep calling it multiple times and it won't delete the previous data.  
                string function = "var win = this.browserbot.getUserWindow();win.errors = win.errors || [];win.errorsJson = win.errorsJson || \"\";win.originalonerror = win.originalonerror || win.onerror ||\"none\";win.onerror = function(errorMsg, url, lineNumber) { win.errors.push({\"errorMsg\": errorMsg || \"\",\"url\": url || \"\",\"lineNumber\": lineNumber || \"\"});if (JSON &&JSON.stringify) win.errorsJson = JSON.stringify(win.errors); if (win.originalonerror != \"none\") win.originalonerror(errorMsg, url, lineNumber);};win.console = {logs: win.console.logs || [],logsJson: win.console.logsJson || \"\",log: function() {win.console.logs.push(arguments);if (JSON && JSON.stringify) win.console.logsJson = JSON.stringify(win.console.logs);},warns: win.console.warns || [],warnsJson: win.console.warnsJson || \"\",warn: function() {win.console.warns.push(arguments); if (JSON && JSON.stringify) win.console.warnsJson = JSON.stringify(win.console.warns);},errors: win.console.errors || [],errorsJson: win.console.errorsJson || \"\",error: function() { win.console.errors.push(arguments);if (JSON && JSON.stringify) win.console.errorsJson = JSON.stringify(win.console.errors);}};";
selenium.GetEval(function);

Now we can get our messages like this:

string errorFunction = "this.browserbot.getUserWindow().console.errorsJson";
string errors = GetEval(errorFunction);
                if(errors!="")
                    TestLog.WriteLine("Console errors found: \r\n" + errors);
       
There you go: Console message logging.  We can of course include other functionality, like throwing an exception if an error occurs, or performing another action.   All we have to do to integrate this into our framework is to perform both actions in our page object base class's constructor, and each page object will automatically call it.  

Monday, March 14, 2011

Config Files

So for this post, I thought I would cover a pretty common problem people have with Selenium....how to get config files working! In my experience, most selenium coders don't have a developer's background, and may not know how to setup a standard App.config file.

When we compile our visual studio project, we end up getting a .dll we open with nUnit.  This is fine, we open the file, run the tests, and yell at the developers just like normal.  However, what if we want to change something about our tests?  Maybe we want to run them against a different environment, or we want to use a different username and password.  We normally would have to go back into visual studio, change the values, recompile, and run again.

However, nUnit automatically looks for a config file with our dll.  We can use this file to store our paramaters we want to change, and not have to recompile to modify our tests.  We could also have several different config files pre-generated, for different environments, users, browsers, languages, or whatever.   The file is saved as xml, so it is editable in any text editor.

Before we go into the file itself you should know that nUnit looks for a config file with your project name + .config.  So if our project is SeleniumTests.dll, it will look in the same directory for a file named SeleniumTests.dll.config.


So lets get started.  To add the config file, in Visual Studio right click on the project and select "Add New Item", and then select Application Configuration File.  Now that we have the file, we need to add our data into it.  So we want to add a new section called appSettings inside the <configuration> tag.  So it will look like this:

<?xml version="1.0" encoding="utf-8" ?>
<configuration>
<appSettings>
</appSettings>
</configuration>

Now, all our variables can be stored inside the appSettings section, as follows:


<?xml version="1.0" encoding="utf-8" ?>
<configuration>
<appSettings>
    <add key="envUrl" value="http://www.google.com/"/>
    <add key="browser" value="*firefox"/>
</appSettings>
</configuration>

We now have two variables created, one with our environment url, and the other with the browser to launch.

The last step is to be able to access our new variables.  The code itself is simple, however I would recommend creating a "Common" class that contains these global variables, that way all of your tests can access them.  In addition, the use of standard get/set commands allows you to use more complicated logic or setup permissions for your variables.

public static class Common
{
   public static int envUrl{ get ConfigurationManager.AppSettings["envUrl"]; set { envUrl= value; } }
 }

So at this point we have a Common.envUrl variable we can get/set.  We're done, right?  Technically yes, but it can be a pain to have to constantly be updating a config file.  Maybe you need to update five variables depending on what environment you are in.  The quickest way to solve this issue is to set sectionGroups.


In our example, we have added the ability to point our tests to different URL's through the Common.envUrl variable.  However, maybe there are other things that need to change as well when we change environments.  Using groups we can specify subsections to switch between without having to edit/delete values.

We define our sectionGroups in the configSections in our config file, then specify the data for each section later.  In the following example, I am defining a timeout value, and my username/password unique for two different environments.

<?xml version="1.0" encoding="utf-8" ?>
<configuration>
  <configSections>
    <sectionGroup name="Environments">
      <section name="qa" type="System.Configuration.AppSettingsSection, System.Configuration, Version=2.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a" />
      <section name="prod" type="System.Configuration.AppSettingsSection, System.Configuration, Version=2.0.0.0, Culture=neutral, PublicKeyToken=b03f5f7f11d50a3a" />
    </sectionGroup>
  </configSections>
<Environments>
    <qa>
      <add key="timeout" value="120000"/>
      <add key="userName" value ="testuser1"/>
      <add key="password" value ="password1"/>
  </qa>
    <prod>
      <add key="timeout" value="60000"/>
      <add key="username" value ="produser1"/>
      <add key="password" value ="password2"/>
  </prod>
  </Environments>
</configuration>

We can now get the section and serialize it as a NameValueCollection. 

private static NameValueCollection envSettings = ConfigurationManager.GetSection("Environments") as NameValueCollection;

Now we can access our variables with the envSettings object like this:

envSettings.username, envSettings.password, envSettings.timeout.

Thursday, February 10, 2011

The Checkbox Quandary

So since I haven't posted in a month, and couldn't think of anything better to discuss, I'm going to share a neat little solution to what I call "The checkbox quandary".  It is summarized as follows:

Clicking on a checkbox only changes it's state.  Selenium doesn't know how to set it to a specific state.  This means that the value of the checkbox could potentially fluctuate each time we run our test.  To get around this we typically check the state of the checkbox before we check it.  Doing this each time is a pain, so here is a simple setCheckbox routine that does the logic for you. It takes standard locator and a boolean specifying if you want the checkbox checked or not.

 public void setCheckbox(string locator, bool isChecked)
        {
            selenium.waitForVisible(locator);
            if((isChecked==true)!=(selenium.IsChecked(locator)))
            {
                    click(locator);
            }

        }

The logic is pretty simple.  If our checkbox state isn't the specified state, we click the checkbox.  Otherwise, it's already in the correct state so we do nothing.  We put a waitForVisible command in there to make sure the element is fully rendered (and all javascript routines on the page are loaded) before we click on it.