Tuesday, May 20, 2014

Software Test World Cup 2014 - Part 2

I wrote recently about the Software Test World Cup. In the last post I said I would post our score card and an analysis if one was possible. First I will give you the raw data, with some "Avg" columns removed. I have not intentionally edited the content other than making the judges more anonymous, format shifting from excel and removing extra columns involving averages and totals. Oh and I marked spelling mistakes with sic, which I just learned should be written with brackets, not parenthesis.

JudgesImportance of Bugs FiledQuality of Bug ReportsNon-Functional Bugs filedWriting/ Quality of Test ReportAccuracy of Test ReportBONUS: Teamwork/Judge interaction 0-10NOTES
A1616914141Bonus: Test Report/Bugs made me want to engage the customer. Several Usability issues.
B15151113131I found that the report did not flow well. I know many teams are expected to give ship/no ship decisions, this [sic] iritates me, i promise not to let it affect my [sic] juding.
C12111014121I like the test report. It gives practical examples, what could have been [sic] testeed for different aspects, e.g. Load, but the ship decision and the "major" bugs seem not to fit imo. bug spread is okay, they tried to consider disability issues (red/green vs. [sic] colorbling ID326)

First of all, I thank the judges for making not only numeric judgments, which is super hard, but also that they spent the time to write some comments. I appreciate that. However, in looking at the judge's comments it seems confusing. For example, C says they liked the report, but their score was no higher than the other reports. They complain that the major issues seem not to fit, such as crashes and the inability to order the product. Granted I don't have the full list of bugs we found, but I wonder what priority they were looking for. That is left to the reader's imagination.

Judges A and B were kinder score wise. Judge A gives us a bonus, but not the bonus we were promised by Matt Heusser, the product owner for answering a question in the youtube channel. Certainly the bonuses were not used how I imagined. I thought our conversations with Matt and the Product Owner would be part of that bonus, but it appears the judges ignored this. I did find it interesting that Matt said that usability would be part of non-functional [no citation, I'm not rewatching 3 hours of video]. We also asked for permission to do some load testing on the Snagit site but never got a response from the product owner, so we chose not to because of legal and ethical implications. It seems like that was not applied in the scoring, but maybe I am wrong. With Snagit as the tool to test there isn't a lot of non-functional testing to do. Judge B on the other hand was harsh in comments but gave good scores (relatively). I agree with Judge B that providing a ship/no ship is annoying, but that is what a conversation is for. We didn't get to have one of those, so we did the best we could.

It is interesting how diverse the opinions of the judges were and also how middle in the road most of their scores were (all items but the bonus were out of 20). Considering how highly we placed, I am guessing either the judges were never impressed or they found out how bad things got and a 10 / 20 really is more like a 15 / 20 relatively speaking. Finally, I promised our test report. Sadly I can't upload the thing to blogspot, so instead I am going to post it below, with some attempts to deal with formatting.


Functional Test Report

By: JCD, Isaac Howard, Wayne Earl, and KRB


Status:

Do not ship

Major Issues:

Undo does not always work, undoing the wrong thing. We found a good number of bugs, including multiple crashes; some seem more realistic seen than others. Bug 241 showed that webcam capture crashed on one particular Mac. When no webcam exists and you attempt to take a camera capture Snagit closes. The Order Now, Tutorial, Get More Stamps and New Output buttons on went to a 404 page. With Preferences open in windows 7, 8, the application refuses to take screenshots. There are about 20 priority 1-3 bugs, which suggests the application isn’t finished.

What Did Work:

We did a performance test with a small Win 8 with 4 gigs of ram and 1.8ghz cpu. It succeeded to take video at a viewable quality. The editor in most cases worked well. Basic usage in most cases works well. The mobile integration worked.

Misc:

We earned bonus points per Matt for giving advice on how to edit video (Jeremy Cd). We asked multiple times in the youtube channel if we could load test the system’s web site but could not get back a response from the judges/Matt. We chose not to ‘hack’ the system due to legal and ethical issues.

Limits of Testing:

- No Automation was generated
- State of Unit Testing is unknown (Customer’s don’t know about unit tests)
- We did even come close to hitting all the menu items.
    o We only have a rough set of tests of windows. Most of our testing was on Macs.
- We didn’t have the technical expertise in the system to capture logs.
- We didn’t have the time to capture the before and after
- We have limited experience with the SUT.
- We only tested the configurations provided. Other configurations of the system were ignored.
- Could not have a conversation with product owner on critical bugs after he left.
- Driver/Hardware testing is limited. We mostly have Maverick OSX.


How Testing Was Planned:

-Website:

As we have been informed that this is a highly hardware dependent product (video and screen capture), which is a commodity, we decided that actually looking at the website of the product might be as important as the product itself. Particularly since a user cannot determine the differences in quality between products without testing the products, the website becomes a primary concern. Particularly since we don’t know they support the types of computers we have, we might be forced to test the site. Finally, if the product has a price, making sure you can’t break the security and get to the download system for free is important.

-Load:

In discussing load testing, we have considered trying to record multiple youtube videos all playing at once, which will stress the hardware and will make for a great deal of variation of the recorder to capture. It might also add a lot of audio channels to record if required. Also using a TV-static screen will push the compression algorithm, as it cannot be compressed well, so we might also test with that. Finally attempt to use a slower system, with little ram and a slow hard disk with lots of CPU and disk usage to see if the recording fails.

-Mobile:

We will attempt to get our relatively few mobile devices to load the system if possible and do some basic usage.

-Usability:

If it is not easy to do the basics of screen/video capture, then users may ask for their money back or go find a different product.

-Feature comparison:
http://en.wikipedia.org/wiki/Comparison_of_screencasting_software#Comparison_by_features
http://lifehacker.com/5839047/five-best-screencasting-or-screen-recording-tools

-Functionality:
• Save a file with a good and a bad file name.
• Record and Capture on app basis, full screen, area
• Two monitors vs One Monitor
• Editing if supported
• Sound if supported
• Pan and zoom if supported
• Arrows, Text, Captions, etc.
• Transitions
• Competitive Intel:
    o http://www.techsmith.com/tutorial-camtasia-8.html
    o http://www.techsmith.com/jing-features.html
    o http://www.telestream.net/screenflow/features.htm
• Long time record (If possible)
• Upload Tools
• Formats supported
• Merging/Dividing recordings
• Tagging / describing the videos other than just the file name
• How easy is it to take the output and use it with another system (e.g. I want to quickly use screen shots and videos from this tool and add them to a bug I created in JIRA)
• Can you add your own voice to the recording? Like narrate what you are doing or what you expect via the microphone on the device you are using
• Can you turn this ability off so you don't hear Isaac swearing at the system or can you remove the swearing track after the fact?

Questions:
- Who are the stakeholders for this testing?
- What are the requirements? What is the minimum viable feature set?
- What is the goal of this testing (Important bugs, release decision, support costs, lawsuits, etc.)?
- Can you please give 3 major user scenarios?
- What is the typical user like? Advanced? Beginner?
- What are the top 3 things the stakeholders care about, such as usability and security.
- Can you give an example of the #1 competitor?
- What sorts of problems do you expect to see?
- Does the SUT require multi-system support? What systems need support (E.G. Mobile, PC, Mac, etc.)? Are there any special features limited to certain browsers/OSes? How many version back are supported?
- What is the purpose of the product? Do you have a vision or mission statement?
- Can you describe the performance profile?
- Is there any documentation that we should review?
- What languages need to be supported?
- Does the application call home (external servers)? Are debug logs available to us? Even if not, is anything sensitive recorded there? Are they stored in a secure location, either locally or externally?
- Do we know what sorts of networks our typical user will use to download the app with? Do we know what the average patch size is?
- Are there other possible configurations that might need to be tested?
- How does the product make money? What is the business case rather than the customer case?
- Should we/can we do any white box testing?
- Can we get access to a developer to ask questions regarding the internals of the system, code coverage, etc.?







Sunday, May 18, 2014

Software Test World Cup 2014 - Part 1

[Part 2 is now posted which includes the scores we received and an analysis.]

I recently participated in the Software Testing World Cup 2014. Just to get it out of the way, our group did place in the semi finals, but did not win. However, that is of less importance to me. Wait you say, what is the point of a contest if not to win. Well for me, I was more interested in getting feedback, learning from the experience and seeing what the idea was. Since this is our blog, not my resume, I'd rather talk about what we did than sell myself. What I am going to try to capture is the experience and outcome of the event.

In starting about the event, I should be clear, a large part of the event happened before the event even started. We were asked to prepare to: Interact with the customer, test the software, write bugs and come out with a report on what was found. We also were told the judging would be on: Being on mission, quality of the bugs, quality of the test report, accuracy, non-functional testing and interaction with the judges.

Knowing that, I did a few things. The first was I assumed it was likely to be web based, as most apps don't work on all platforms, so I started to enumerate what sorts of non-functional tests we could do. This turned out to be wrong, but it is better to be over prepared than not in my view. Then I looked into how to write such as report as I don't do those sorts of consultant-oriented paperwork. I have conversations with the developers and work closely with the product owner. Once I had an idea for the format, I created a list of questions we would want to ask, some of which were perhaps excessive, but again, I assumed more was better because in the moment I could trim the list by the way the customer answered questions, but I might fail to think of a question which could not be 'included' unless I spent time thinking about it. Ironically this felt a little heavy (for a group of judges which is composed of lots of pro CDT people), but again, these sorts of reports are heavy in my view.

At this point we had generically pre-processed the report and the interaction with the judges. We were told the day of that it was visual screen capture software so I did some research before the start of the contest, to compare feature sets. I captured some possible tests we might do and documented those.

While I have built my own screen recording tools, and even automation tools I have rarely gotten to the chance to test out such a massively well known support tool as SnagIt. We didn't know it was SnagIt until half an hour before the competition had actually started. With half an hour to go, we started to organize our thoughts and take a look around the tool. I create a few product oriented questions and found some obvious bugs. One the really concerned me was the order page link did not work. That turned out to be important later. I wrote notes down, but didn't file bugs as the contest had not started. I was just trying to understand the product.

We were told we could start asking questions, and I started peppering the judge with questions, but try to keep them slow enough to know if I need a follow up question or not. The judge's ability to respond to feedback was a little poor. For example I wrote a question regarding screen capture in Firefox which includes a plugin, something that seemed poorly documented. I missed saying it was for Mac as the plugin is only for Mac. They got confused by this and I updated my response but they were not looking at the chat log. I don't mind the video chat, but that became frustrating for me.

We as group decided two people would be in Mac land, one would be in windows VM and I would be between my personal win 8 box and my Mac.  We split up the functionality as well and started to test.  We did a load test using static fuzzing for a recording (E.G. something that wouldn't compress well) and my win 8 box has very little resources, so I did a load test where I recorded video while compiling.  I know lots of other testing was done, but those were some of the more interesting tests that I can recall (this was all written weeks after the contest was over).

As a tool smith who specialized in Windows, I appreciate how hard it is to build complex tools. However, when the product owner said that they used the tool to capture screenshots of bugs in the tool it seemed funny as we wrote up several bugs around that. We found that claim to be incorrect, but I wonder if the judges heard/recalled that claim since I didn't create strong paper-trail around it. Also because the judges in general don't have conversations with people afterwards it is hard to know what someone is thinking. I get that writing has to be clear, but how much context does one normally capture? Without knowing the culture, that might have been a throw away line from the product owner or it might have been serious. It makes me appreciate consultants in some ways, even if I do think they sometimes are more about conning and insulting than providing useful data.

It was genuinely neat to try to test something I would not have tested otherwise and doubt I will ever personally work on. It was fun. We found a lot of bugs. With one hour left I asked for permission to load test the download link, as that seemed part of the system and non-functional. We couldn't get a response. The judges were joking and talking about their experiences. Isaac tweeted two of them and they ignored them. The judges were filling the dead space that should have been silent with conversation, which I think was a detriment to the experience, but on the other hand, maybe that is part of the contest. Simulating that annoying coworker who keeps talking and talking in one long unbroken sentence moving from topic to topic... Yeah, Wayne! Quit that.

While finalizing our report, the judges sent a response to one of my early findings about the order link being broken saying that was expected.  I had put in a post 2+ hours earlier about that bug and the judges had not answered.  I didn't find out they had responded until after the contest was over.  We were given a chance to amend our report because of this, which I appreciate, but I also think it might have changed our testing strategy.  When you find one bug, you often look for clusters of bugs, which is what I did.  Obviously life is not fair and I don't expect the first year of a contest to run perfectly smoothly.

At this point we know our placing but not our scores. The contest was broken up by judge with points and scores attached. We got bonus points for helping one user, or at least that is what we were told. I wonder if we will get them or if their will be fear that giving said scores might cause disputes like you forgot to give points for X. It would suck if our placement went up or down at this point. But we shall see.  I will post a second blog entry based upon the score information we are given.

Friday, April 18, 2014

A Personal Failure Analysis: Why it appears CDT has failed

Talks & Failure Analysis

For the last few weeks I have been doing one presentation a week to various monthly groups and to a yearly event.  I have one more planned presentation in June, but this really isn't about my presentations, but rather the audience and fellow speakers.  

The first presentation was a meetup group Isaac and I helped form up in March.  Isaac and I gave the talk and it went really well.  Our audience was maybe 40% high-level QA with a few junior and the rest mid levels.  The talk was about testing an object in black box style.  We gave this talk 3 times total, and each time it seemed to go a little worse.  

When we did it with a group of developers, we moved to more white box/boundary testing style.  We had one person refuse to test out of principle.  In part I think us changing the style was at fault, but two sharper developers really got it, so I don't think it was totally our fault.  Two people really didn't get it until I changed boundary testing into a UI exercise.  This tells me even (junior) developers have a hard time visualizing what code does in their head.  

The final audience was a mixture of testers and developers.  The developers were often dis-interested in our talk and several of them left to go see a different track.  Either they felt applying black box techniques was not part of their domain or that they already knew this and wanted something new.  The irony of that is that 60% of the sessions were marked beginner, so learning something 'new' for a mid/senior developer is challenging at best.

I had a second talk which was a shorter version of my intended lecture I will be giving out of the country.  It was around reflections, something I think I have some level of mastery around, so I felt very prepared for my talk.  I will admit my cadence was a bit off, I probably rushed a little too much and perhaps didn't give enough audience breaks, but it was a code oriented talk.  Given the complexity, that might have been a small part of the problem but in reality, most of the audience were either junior developers or students who in their 2nd-4th year of college.  I had marked my talk as advanced but I couldn't control the people who attended.  This caused me to wonder something.  Why is it we are still talking about TDD, Intro to jquery, etc.?  I don't mean to say we shouldn't teach new people, but why are we mostly limited to beginners?

Why do we retrain people all the time?

In talking with 'experts' in the business, it becomes clear that my question of why do we repeat ourselves is because in some communities, such as testing, people don't stick with it in the long run.  Roughly 5 years appears to be the mean time between starting a career in testing and going to something else.  The second issue is that many people are not particularly interested in expanding their knowledge.  They just want to go to work, do their job and go home.  Don't get me wrong, I don't mean to say you should work 90 hour weeks, but many people have little passion for their work or improving their craft.  

Sometimes managers make these engineers look into 'new' development techniques, which create such a long tail into training that something like TDD continues to need talks 10 years later.  

With that analysis, I want to talk about it with an 'expert', and so I found the keynote speaker whom I spoke to at length.  I don't want to cite his name as I have not asked his permission nor do I have an easily accessible method to communicate with him (I don't have his email address and don't use twitter).  

He is a well-known developer who has a pod cast program with hundreds of different interviews with various people within the software industry.  In our discussion, I asked him if he knew Matt Heusser and he said he didn't.  I asked if he knew of James Bach, Cem Kaner, Rex Black...  He knew of no one I asked of, although having reviewed his interviews, the only tester I recognized that he interviewed was James Whittaker, which was many years ago.  I had not asked if he knew Whittaker.  I did ask if he knew any US experts/thinkers in test, but he declined to answer my question, saying he didn't think that way.  Perhaps my question was somehow confusing.

In speaking with this gentleman, a man who clearly had a wide variety of experience, and seemed to subscribe to the Analytical School of Testing, it became clear to me.  One of the reasons we end up not keeping testers around of long is because we as a group are not well known.  James Bach, in spite or perhaps because of his controversy, has one of the biggest names in testing.  Sadly, managers aren't demanding people start doing 'Bachian testing' (even though I suspect Bach himself would object to that name and concept).  We don't have a glamorous single one size fits all solution, like Scrum or TDD.  CDT is so vague no one but its practitioners knows exactly what it is.  It isn't easily measurable like Scrum or TDD which I can say if we are doing it or not (Are there sprints?  Are there unit tests?).  It makes me wonder if CDT has lost its war when managers can't easily measure if there CDT testers are even doing CDT testing, much less if CDT is working for them.  The metrics (be it sprints or number of unit tests) might ultimately not matter, but managers still at least one security blanket before they are likely to invest in structural change.

Solutions?

I'm not sure I have any useful solutions other than to say this appears to be a community issue.  I have tried doing mentor-ships with mild success.  It however doesn't scale beyond 1-2 people at a time.  I have tried doing lectures, but as soon as I stray past beginner style thinking or just talking about the concepts without talking about implementation, it seems to be clunky at best.  I have started a group with Isaac to see if a more personal monthly audience will help.  Maybe there we can make headway, but I don't know.  One nice thing with it being monthly and having multiple speakers, I might find out if the issue is in fact somehow with me and how I am presenting.  Even outside of presenting, maybe the way I communicate poorly and that just needs improvement?  For that matter, maybe I'm using the wrong language (English) to reach my audience of highly-motivate testers and software developers?  Perhaps there are more motivated testers outside of the US and my issues only revolve around where I live?

Perhaps I need to accept that I can only speak to a handful of experts, but how do I find and group them together?  How can the experts then package this stuff up together in such a way that they might learn from my knowledge?

Outside of myself, I know Kaner is moving to a more teaching oriented approach generating classes in which the students come to him.  Having taken a BBST course and talking to others who have, I know only about 1 in 3 (or less) of those students seem to both try and have the skills to succeed.  Some choose only to pass by the skin of their teeth and others simply don't get it.  Maybe some of it is because the course is in English, but it certainly isn't the only reason.  However, 1 in 3 might be as good as it ever gets, perhaps my expectations are just too high.  Kaner appears to have given up on a education only system by adding in certifications.  It appears to me the ultimate answer maybe that one has to market one's ideas more in order to convince people to change.  Perhaps the reason Kaner is talking about moving to such a path is because he failed to market his solution well enough.  Going back the keynote speaker I spoke to after my talk, he suggested that if he didn't know about the thinkers/experts of test or the ideas of CDT, they must not be that important.  Because if you don't hear about it, it can't matter.

Thursday, March 20, 2014

What is a comment? A miserable little pile of words!

A few weeks ago I tried an experiment.  I tried to comment on at least one test-related article or comment every day of the week.  I tried this for about a week and a half.  It was tiring, but interesting, with a result that made me sad.  I didn't do wimpy non-comments like "Great job".  I did in depth, well considered comments, that often took me an hour to develop.  And I got nothing back.  It gives me a new respect for people like James Bach who seem to be all over the Internet commenting all the time and doesn't seem to feel defeated.  As a consultant and trainer, that is a method of advertisement, but I do believe James Bach genuinely cares too.  However, my concerns with comments were developed much earlier...

I am a big fan of  Code Horror by Jeff Atwood who influenced my choice to do a blog in the first place.  While I don't always agree with him, I love his views on commenting.  In fact, he and James had a 'chat' about comments some years ago.  To summarize, James Bach had noticed a major flaw in Jeff Atwood's blog post and wrote his own post on it.  Jeff wanted to continue the conversation but couldn't because James didn't allow for comments.  Now that was years ago and once again, to Bach's credit, comments are enabled.  Victory.  We are all done.  Except...

Cranky Old Man


I still have complaints about how our community uses comments.  Certainly we have useless 'great job' comments without any critical thinking involved.  You can see them in even some of the bigger venues.  Even if you think the blog post is great, and you're their new biggest fan, that sort of comment helps no one.  I won't think about you, my biggest fan.  If you hate my work and yell at me without any particular logic, you're no better than the old man to the left.  However, as a blog writer, am I just a cranky old man yelling at human nature?

I don't think so.  Often Slashdot's comments are valuable and well thought out.  Bach's comments section is curated, often insightful and from some reports perhaps even censoring, but that is a different issue.  I think there are a few issues with our community.  The first is that we write too much.  You heard me.  An author says we write too much.  Ha!  I can read about 5x more than I can write.  When Markus Gartner on a nearly daily basis writes a 1-2 page blog, I think it is nearly impossible to have a conversation about any topic he brings up.  In two days Markus is off into the next subject.  Being 12 hours away doesn't help the conversation, but in the world of twitter, where detail doesn't matter but lip service is paid, you can't develop consensus or new ideas.

Why can't we slow down and as a community genuinely talk through issues?  I am amazed at those whom write daily, and I get that people like Matt Heusser who suggest writing at least 2000 words twice a week.  It isn't an invalid idea for improving writing, but that doesn't mean everything you write is worth posting.  I currently have some 2120 [this article is now published!] articles in draft format, including this one.  Not all of them make it out to you, the public.  I want to give you quality work, and hope to get quality comments back.  I don't write too much, in the hopes that I can learn from the community as I write articles.

But Wait...

What about all those spammers....?

Well to be honest, I rarely see them here.  I think Jeff Atwood might have a solution for bigger blogs, but I am reserving judgment for now.  Truthfully, that would be a good problem to have, as that would mean there are a lot more people caring about QA. 

What if I don't have anything to say?

Really?  You read a 1-2 page article, and found nothing of value to say? Nothing was added to your mental model?  You didn't see any mistakes, any logical leaps too far nor anything understated?  The blog post was so dull it would have been better off being ignored?  Which leads me to...

What if what I have to say is too mean?

Well, for this blogger, I give you permission to write mean things about my, JCD's blog posts.  In fact I WANT you to.  Please, say you hate it and then justify that opinion with facts.

...and if you do write a comment, with insight, which pushes me to write a response, I will try very hard to reply if not write an entire additional blog post on the subject.  So please do!

* Title with apologies to André Malraux.

Monday, February 24, 2014

Look, I wrote something.

Well I haven't been writing much as of late.
However I haven't been slacking...

http://www.stickyminds.com/article/helpful-tips-hiring-better-testers

I've just been focusing on a different venue for my ramblings.

Tuesday, February 11, 2014

Being a Fraud and a Failure

Often I feel like a fraud.  This comes in multiple sizes and shapes and I think it is often ignored in creative industries such as software because we liken ourselves more to scientists rather than artists.  Artists already know the feeling I am thinking off.  Art & Copy had some of the best Marketer's saying things like:
The frightening and most difficult thing about being what somebody calls a creative person is that you have absolutely no idea where any of your thoughts come from, really. And especially, you don't have any idea about where they're going to come from tomorrow. - Hal Riney
and
I think most of the creative people are so damn insecure that they want to think they know everything, but they know deep in their hearts they're just in deep trouble from the minute they get up in the morning. So if you can tell them "that's what you're supposed to be", that's kind of liberating. - Dan Wieden
When I go to test something, I start out with panic for a second or two, wondering what the hell I am going to do.  Then I draw from my experience and start developing tests.  But that panic for someone with less experience might last a good deal longer.  In CDT, there is a sort of freedom given, because 'it depends' on what your context requires.  I sometimes wonder if it isn't that there is a science, a way of telling what is best to do, not because context gets in the way but because what testing requires is as much art as science.  I may not be a pretend expert, but I know much of my creativity in testing comes from some place that is not just from my smarts.  Sure I have science to back up some of my testing, but I also have 'hunches.'

Hunches are not always easy to swallow when I know they could be wrong.  If I only have time for one feature but two features need testing, I could test feature X on a hunch and fail the entire company when feature Y is broken in a way I had not predicted.  Elizabeth Gilbert suggests we simply shift blame from the personal 'I failed' to a third party entity such as a muse, in which a muse hinted to me, giving me a hunch.  I don't find this a particularly satisfactory way of dealing with my feelings of being a fraud, failing at my job.  I don't like putting credit and blame onto a external entity, so I need another way of dealing with my feelings.

Alternatively, I can read testing blogs where we all collectively struggle to find better ways to improve our chances for success.  This can go all the way into the more scientific side to testing.  If only we had enough data, we could kill off the hunch completely.  The Factory School seems to see it as a problem of Taylorism, if only we could just lock down all the changes.  Document everything, know everything, then hunches aren't needed.

The more I research, the more I think our industry is full of fraudulent feelings with different people attempting to resolve them by adopting 'proven' systems and techniques.  They aren't the fraud, considering their sincere attempts to learn, but they have to find a way to limit failing.   So they turn to these 'experts' to solve your problems.  They have the answers, and you should just trust them.  They provide systems and will teach them to you, often for a fee.  However, if you fail, you aren't doing their system right, not that their system is broken.  You are forced to feel like you failed to heed the words taught, and only if you had worked harder, followed the rules closer, etc. you'd have succeeded.  These tools make us feel the failure and not that the tool failed.  Who is to blame, particularly if the tool often succeeds?

Then we move to team oriented solutions such as Agile.  Agile says it is a team effort, so if a bug is found in production by customers, the team is to blame.  Moving from an "I" to a "We", much like the muse.  This is a paradigm shift, even if not a perfect one.  You now can know if you are a fraud or not based upon the feedback from the team.  The problem is you cannot know if the team is filled with frauds, if the team is just being polite or if they are simply more experienced then you are.  To assume the group will be wiser than yourself, and able to judge your success has its limits based upon who they are.  To assume you are able to judge the team also maybe flawed.  To quote Robert Heinlein,
Democracy is based on the assumption that a million men are wiser than one man. How's that again? I missed something. Autocracy is based on the assumption that one man is wiser than a million men. Let's play that over again, too. Who decides?
Every method is flawed.  I personally like the method Art & Copy suggests.  Rather than blaming a third party or putting your faith in someone else's system, I think it is healthy knowing that sometimes you will fail.  That the feelings of being a fraud comes, but with as many wins, as many bugs as you've found, you aren't just pretending to test, you aren't a fraud.  You really are helping the customers, finding bugs early on, helping resolve issues.  Maybe you do fail, but pick yourself back up and do it again.  However, that is personal preference.  I don't think there is an answer.  Some will like science.  Some will like the sales pitches.  Some will like the art.  Some will like the team.  Some like having no right answer.  They are all fine and all flawed in their own way.

Thursday, January 30, 2014

Book Consideration: Rethinking Systems Analysis & Design

To be clear, this is my second book by Gerald M. Weinberg, and I’m not reading the books he wrote in the particular order he published them in.  I like the author’s style of thinking, but in the Introduction to Systems Thinking book, he was very general in his descriptions, hitting a great deal of subjects, with his thoughts.  This of course is a big part of system’s thinking – the idea that you can apply what you know in one field against many different fields using a form of logical thinking and general rules.  While some of the design of the book could have used some refactoring, it was a good start.

In Rethinking Systems, the book focuses mostly on the development of software systems and how this thinking can be directly applied to software.  While some of the ideas are fairly standard (and the book was published more than 10 years ago), it provides some interesting insights.  He starts out by considering analysis vs slogans.  He talks of how we sometimes we over use history as a method for predicting the future.  While history is valuable, analysis will provide data not found in history.  That is to say if X is good, X+1 might not be better, depending upon the attributes that we want out of X+1.  However, we might have a slogan “Bigger is better [for X]”, which might be true up to a given point, where the system then fails to scale for some reason.  He warns of how software developers are particularly vulnerable to certain methodologies, because [in my reading of his opinion] we as an industry don’t have a long history, and even a 5% increase in productivity is great if we consider the number of failed projects in the past.  Unlike the physical world where you have to build a new bridge every time, in spite of it being very similar design, where as we in software always build something new, else we could just copy the previous build for “free”.

Chapter two talks of what actually makes up a [software] system, with a fairly reasonable picture, including the external world, the organization, training materials, training done on the job, wetware, existing files, test data, job control, (and all of this before), the program and hardware (Pg 34).  I thought it was a good point that part of the investment in software is the people and the wetware (the knowing what code is where because you built it which is faster than reading code or documentation).  He also talks about how we need to use our history more wisely, as we don’t often understand why something is the way it is.  While not an example given, I would say the Unicode set(s), is a good example, as they are design for different forms of optimization, but are fairly ironic since they were in fact an attempt to solve a very real world issue ASCII had, as computers were originally designed for English.  I also think of Joel’s amazing article on Martian Headsets (It is worth reading.  Go now, I’ll wait).

Chapter 3 is about how the observer fails often to observe anything of value (including thinking they saw X but saw Y), or at least the wrongs things (like how magicians have you look the wrong way).  In QA this is an important point, and one which I have suffered through more than once.  I actually talked on this a little in a presentation I gave about changing one’s perspective, a valuable thing to do.  He talks about studying the existing system to understand rather than criticize.  This is an important lesson, I hope to take more to heart, as a QA engineer, I have to be careful not to be hurtful to those who created the code, but more of an impartial observer, stating what I saw and why it seems wrong, not just to me, but from a more broad perspective.  In some ways this reminds me of a fair witness.

Chapter 4 talks of self-validating questions, that is to say, questions that require a response that will in fact validate an understanding of the question itself.  He talks of the question “Does that contain special characters?” which gets the answer “no”.  He assumes special characters means alpha + numeric, but nothing else.  In this day and age, it would mean something else, but by asking a yes/no question, without a follow up can get you into trouble.  He speaks of the problem that programmers tend to be defensive in what they do and how they act because users accuse them of deliberately causing trouble if the program says “don’t do X” but the user demands it anyway.  On the other hand, if they just ignore the user and do what they want, they can stubbornly miss important details.  This is a trick problem I also talked about in my talk with dealing with users.  I felt he had good insights and this chapter alone was worth the effort.

The last few chapters are on design of software, including the philosophy, trade offs and the mind of a designer.  I really don’t have much to say about these chapters.  It is not that they are bad, but it was not the most exciting to me.  He talks of being aware of your designing, not to over design nor to ignore the reality of the world.  I find these not exactly contradictory, but certainly a narrow path to follow.  To me, he could have done an entire book on design, but I suspect his strong point is system analysis, not design.  He suggests designing for understanding (know your audience), strike balance between variation and selection (don’t be too original or ground breaking for your audience), etc.  He speaks of trade offs and how to represent them as curves, which is a somewhat novel way of displaying the information, but fairly obvious to anyone who knows that famous triangle with fast, good, cost (pick two).

There are three other things which I would like to mention, but which I have not found (in my consideration of the book) what chapter they were in:

  1. He speaks of spell check which is ‘obviously good’, but is it?  It seems to him that the cost of the grievous errors might be greater than the minor typos.  He noted examples where this caused confusion as he had intentional mistakes that were fixed by editors.  He actually complains about one editor using glue and razors to fix an intentional bug in a different book, which I found pretty funny.
  2. He speaks of students in his class using systems analysis to gain ‘masters level’ knowledge in a subject and that on average some subjects took just 3 months, but others took longer (he didn’t say how long).  He specifically said English was a subject that took longer and computer science was one of the shorter subjects time wise.  I have my doubts, but interesting none the less, particularly when I recall an (apocryphal?) story of a man who said he could pass any test for his degree.  He was studying something like psychology, but said he could pass any test and thus was given a test in dentistry, which he got a B in, getting his degree.
  3. I found it interesting that in the author’s epilogue, he states: “basic human needs… air, water, food, sex… the need to judge other people.”  What I find interesting about it is I have had several talks about this subject with various people, one of whom claims not to judge, when judging is defined as “A person able or qualified to give an opinion on something.”  The author continues, “To write a book… you have to have an uncontrollable urge to snoop and pass judgment… to study what people do and tell them how to redesign their activities.”  It seems to me that judgment is of utmost importance in both the high and the low.  In the high level, we build something we judge to be useful to others, in spite of never having perfect information (making one not qualified?).  This egomaniacal belief that we are good enough to play a sort of demigod, handing design from on high, with the assumption that what we do will be good, or at least better than having not.  We design not truly ever knowing what our users really want and at best we can develop heuristics on this, but not hard and fast rules.  On the low, we design the system’s structures with limitations based upon our best understanding of what will affect it, judging how the system will be used (E.G. Computers will only need English characters) and who will use them, not really knowing what the future will bring.  The trick in my opinion is to be wise enough, not to intend harm with our judgments, nor be too harsh; for often times we end up blinded by the judgment, unable to let go based upon our assumptions.  Judgments, it seems to me (based upon the author’s words) are just another heuristic, not a truth.  In my own view, judgments are a non-binding, unenforceable guesses (with a probability weight behind it) as to how something works, based upon previously observed factors.
In my opinion, I would read the systems thinking book first and then read this book as Systems thinking gives you a broader base of theory, but this book is more practical and does have value.

Thursday, January 16, 2014

Why can't anyone talk about frameworks?

In writing for WHOSE, I was dismayed at the total lack of valuable information regarding automation frameworks and developing them.  I could find some work on the frameworks with names (data driven, model driven and keyword driven), but almost nothing on how to design a framework.  I get that few people can claim to have written 5-10 frameworks like I have, but why is it we are stuck with only these 3 types of frameworks?

Let me define my terms a little (I feel like a word of the week might show up sometime soon for this).  An architecture is a concept, the boxes you write on a board that are connected by lines, the UML diagram or the concepts locked in someone's head.  Architecture never exists outside of the stuff of designs and isn't tied to anything, like a particular tool.  Frameworks on the other hand have real stuff behind them.  They have code, they do things.  They still aren't the tests, but they are the pieces that assist the test and are called by the test.  A test results datastore is framework, a file reading utility is framework, but the test along with its steps is not part of the framework.

Now let me talk about a few framework ideas I have had for the past 10 years.  Some of them are old and some are relatively recent.  I am going to pull from some of my presentations of old, but the ideas have at least been useful for one framework of mine, if not more.

Magic-Words


I'm sure I'm not the first one to come to this realization, but I have found no records of other automation engineers speaking of this before me.  I have heard the term DSL (Domain Specific Language) which I think is generally too tied to Keyword-driven testing, but a close and reasonable label.  The concept is to use the compiler and auto complete to assist in your writing of the framework.  Some people like the keyword driven frameworks, but in my past experience, they don't give compile time checking nor do they help you via auto complete.  So I write code using a few magic words.  Example: Test.Steps.*, UI.Page.*, DBTest.Data, etc.  These few words are all organizational and allow for a new user to 'discover' the functionality of the automation.  It also forces your automation to separate out the testing from the framework.  A simple example of that can be given:

@Test()
public void aTestOfGoogleSearch() {
 Test.Browser.OpenBrowser("www.google.com");
 Test.Steps.GoogleHome.Search("test");
 Test.Steps.GoogleSearch.VerifySearch("test");
}

//Example of how Test might work in C#, in Java it would have to be a method.
public class TestBase { //All tests inherit this
  private TestFramework test = new TestFramework();
  public TestFramework Test { get { return test; } }
}

Clearly the steps are somewhere else while the test is local to what you can see.  The "Test.*" provides access to all the functionality and is the key to discoverability.

Reflection-Oriented Data Generation


I have spoken of reflections a lot, and I think reflections are a wonderful tool for solving data-generation style problems.  Using annotations/attributes to tell each piece of data how to generate, what sorts of expectations there are (success, failure with exception x, etc.), filter the values you allow to generate and then picking a value and testing with it is great.  I have a talk later this year where I will go in depth on the subject and I hope to have a solid code example to show.  I will certainly post that up when I have it, but for now I will hold off on that.

...

Okay, fine, I'll give you a little preview of what it would look like (using Java):

public class Address {

 @FieldData(classes=NameGenerator.class)
 private String Name;
 @FieldData(classes=StateGenerator.class)
 private String State;
 //...

}
public class NameGenerator {

  public List<Data> Generate() {
   List<Data> d = new ArrayList<Data>();
   d.add(new Data("Joe", TestDetails.Positive);
   d.add(new Data(RandomString.Unicode(10),  {TestDetails.Unicode, TestDetails.Negative));//Assume we don't support Unicode, shame on us.
   //TODO More test data to be added
   return d;
  }

}

Details


Why is it that we as engineers who love the details fail to talk about them?  I get that we have time limits and I don't want to write a book for every blog post, but rarely do I see anyone outside of James McCaffrey and sometimes Doug Hoffman talk on the details.  Even if you don't have a framework, or a huge set of code, why can't you talk about your minor innovations?  I come up with new and awesome ideas once in a while, but I come up with lots of little innovations all the time.

Let me give one example and maybe that will get your brain thinking.  Maybe you'll write a little blog on the idea and even link to it in the comments.  I once helped write a framework piece with my awesome co-author, Jeremy Reeder, to figure out the most likely reason a test would fail.  How?

Well we took all the attributes we knew, mostly via reflections of the test and put them into a big bag.  We knew all the words used in the test name, all the parameters passed in, the failures in the test, etc.  We would look at all the failing tests and see which ones had similar attributes.  Then we looked at the passing tests and looked to see which pieces of evidence could 'disprove' the likeliness of a cause.

For example, say 10 tests failed.  All 10 involving a Brazilian page. 7 of those touched checkout and 5 of those ordered an item.  We would assume that the Brazilian language is the flaw if all tests failed, as that might be the most common issue.  However, if we had passing tests involving Brazilian, then that seems less likely, so we would see if we could at least establish if all checkout failures had no passing tests involving checkout.  If none had, we would say there was a good chance that checkout was broken and notify manual testers to investigate that part of the system first.  It worked really well and solved a lot of bugs quickly.

I do admit I am skipping some of the details in this example, like we did consider variables in concert, like Brazilian tests that involved checkout might be considered together rather than just as separate variables, but I hope this is enough that if you wanted to you could build your own solution.

Now your turn.  Talk about your framework triumphs.  Blog about them and if you want to put a link in the comments.

Friday, January 3, 2014

Heroes and Villains

Heroes

I have a host of people that I read and read and read.  These are my literary Heroes.  They constantly give me new gems and bright insights.  I pale in my work before these folks.  The sad news is that my heroes have slowly slipped away and all I have is their tremendous bodies of work.  In the blog-arena, I really appreciate Jeff Atwood, Joel Spolsky and Steve Yegge (who has multiple blogs).  Just as an example of my reading, I am all the back to 2005 in Jeff's blog, reading it backwards entry by entry.

Now two things you'll note from that.  One, I really enjoy what developers have to say and I really don't have a lot of test-related bloggers I hit on a regular basis.  Even if I start widening my knowledge net, Martin Fowler, Scott Hanselman and those Dot Net Rocks podcast guys probably has beaten out other test-related heroes I have read.  Now I do tend to focus on what they have to say on test when they have something on test, but I spend way more time on the development process than on capital-T Test.  These guys are brilliant and have wonderful things to say.  They come up with all sorts of clever coding patterns and practices.  They consider the process as a whole and care about the craft.  They are Heroes.

Then I should consider my list of more generic Heroes.  Heroes of science like Richard Feynman, Heroes of science fiction like Robert Heinlein.  Heroes of thought like Rene Des Cartes, Heroes of humanity like George Bernard Shaw.  These are all men I have read and found to have enlightening things to say, although some are rather obscure.

Right now, this is how I feel about test and "Heroes".  There are no Heroes.  There are a few knowledgeable people, but the fractured nature makes it hard to pin down.  Most people who come into test come into it by 'accident'.  We are still forging paths and hitting dead ends.  Yet we do have another aspect.

Villains

Perhaps our lack of Heroes is the nature of our business, so at best we get Villains.  Well, Villains are cool, right?  Who doesn't like a good Super Villain?  The Riddler or Magneto come to mind.  They have super powers, they do cool things and in the end, they often feel justified in their deeds.  I hesitate to call anyone a "Villain" in a community I work in.  Worse yet, CDT might be so accepting that Villains are considered good guys in their own way.  Now if I was James Bach, I would have "“depraved” enemies".  Now does that make someone like Rex Black a Villain in Bach's eyes?  Is the visa-versa true for Black?  Am I a sort of Villain of the CDT community for asking questions?  In talking with one senior-level tester, I could be seen as a Villain.  The all-mighty dollar of consultancy demands for true purity of our testing Scotsmen and my questions for CDT is dangerous for those dollars. I won't answer these questions regarding villainy for you, but I do want to be clear, I am not calling any of these gentlemen Villains.

The downfall of our method is that we don't have any golden boys of testing.  Instead, we challenge each other to get better but don't have an easy way of showing off our skills.  I can't even personally say if Kaner, Bach or Black are good testers.  I can look at some of Atwood, Spolsky and Yegges code and judge them, as that is often visible.  Testing is often an invisible task, not one with an output easily examined.  I think some of the information of the test community can be useful, but I don't use all of any practitioner's 'methods'.

I am left with two things.  Are there any real Villains of testing and can we have Heroes in testing?  How can I evaluate a given person's skill in "Test" compare to writing or management or bug writing or all the other skills that make up "Testing"?  Even with testing an application, it often requires parameters that are hard to control, like the versions of software, builds, time, memory, threading and all the other 'impossibility of complete testing' pieces.  So even with bug reports, test cases, etc. it might be impossible for me to truly evaluate the output.  If I can't know (there are plenty of heuristics, but actually knowing is much harder) someone is a good tester, then I guess I'm just left with is a Villain just a perspective or are their absolutes?

I don't know that I can answer that, but maybe I'll look into getting an outfit, just in case.

Thursday, January 2, 2014

Exploratory Software Testing: Tips, Tricks, Tours, and Techniques to Guide Test Design

Exploratory testing, according to James A. Whittaker, is “When the scripts are removed entirely…, the process is called exploratory testing.”  Whittaker’s statement taken at face value assumes that you cannot perform exploratory testing with a script, although he does later revise this to include what I might call “partial scripting” (running a test up to a point and then exploring from there) and Whittaker calls exploratory testing in the small.  He also looks at exploratory testing in the large, which is what I typically consider exploratory testing (which I will cover later on).  In my opinion, the problem with his initial statement and exploratory testing in the small is he is actually describing testing various states just without written steps.  Really, you could easily create a simple table of states and outcomes and it would be just scripted testing.  This to me is not exploratory testing.  To be fair, he later on acknowledges that you can in fact combined scripting and exploratory testing, but that seems to not be covered in his definition of exploratory testing (which I find ironic since it is his primary purpose for the book and as a tester he should have noticed this incongruity).

Exploratory testing in the large, as I described earlier is the act of testing one or more areas by using non-scripted patterns.  Whittaker describes a system of patterns, which he calls Tours.  Tours are, according to Whittaker, a “…mix of structure and freedom.…”  His Tours are designed around either attempts to change your point of view, the methods of testing or the scope of the testing.  Personally, I find the names of the Tours as very poor representatives of the concepts he is trying to put across.  For example, the Money Tour could have easily been called Customer Demo Testing or the FedEx Tour (one of the better name Tours) could be called Data Flow Testing.  I suspect that the reason for the funky naming convention was either so he could use the word Tour (and thus the analogies) or the so that he could make a book rather than a long blog.  While I might object to his implementation, I can’t object to his goal of making better testers.

My other objection comes from the limits of what he ultimately came up with.  He describes to a very limited degree the attempt to see the application from other user’s prospective.  He tells stories of tourists, but I didn’t notice him explicitly link the tourists to users, what the users do and what they expect.  He uses these tourists to ask questions, but not to change one’s bias from one type of imagined user to another.  For me, when testing, I make a great deal of effort to simulate various types of users and to intuit what a given user would do.  If I see a button, I might see 6 different styles of users.  The elderly person who doesn’t get it is a button because it doesn’t look like a button, the expert who wants to middle click the button to open in a new window, etc.  Even then I sometimes go further and imagine not only the users but what their reactions might be to a given system.  What it is that they would want, where they would get confused, where they want the system simpler rather than more complex, etc.?  The book does not seem to address this type of cross-user non-functional style thinking (exploratory testing?) or what type of outcomes might be generated from it.


Outside of the actual techniques addressed in this book, another question left unaddressed is how one manages the various forms of testing in harmony.  I realize this is a complex question, but that is why one pays good money for the book.  To be clear, I don’t mean, what Tours to choose for a given project, which he does address through examples (ironically not written by the author; how many pages did he actually write?).  I do mean, when Tours, when scripting, when automating, or when doing whatever other types of testing the author might have cared to mention.  


NOTE: In this post, I use the term tester as a generic term including Test Engineers, SDETs, Developers writing unit tests, hackers attempting to break in or even customers on a beta product.  In all cases, the person is testing the system.  Only customers on production systems should not be considered testers, since at that point the product should have a fairly low number of bugs, and the customer is not looking for defects.



Unstructured Additional Notes [I wanted to include some unstructured misc thoughts on little bits and bobs that I wanted to make some minor comments.]:
  • On page 120, Memorylessness is described, which I think is a huge problem with testing outside of monolithic organizations, but not exactly in the way described.  The problem he describes includes testers forgetting what they tested (a reasonable problem) and testers forgetting what testing techniques find large sets of bugs (a mostly unreasonable problem IMHO and too easily abused by management).
    • He should also mention how it is hard to “remember” what is automated already, as that too seems to be a serious issue.  I suppose perhaps his “Testipdedia” is the closest thing he had to a solution, but that seems like a very manual effort or so generic to be of little value.
  • The WMP and VSTS (Page 97-111) testing comments were of some value.
  • The idea of tracking bugs by % of effort (time wise) to % of bugs found seemed like a good idea (Page 104).  Also the visual test tool falls under a similar area, where testers are able to see where areas have the most functionality, testing/automated focus, bugs, etc.
  • I think his commandments where worth reading, even if only 5 pages of the book.
  • I found his comments on measuring a tester by the improvements to the development team interesting, but rather hard to measure.  Do you look at bug trends, points completed by the team or what?  To add to that, I think that quality of the design (both interface and code) can be part of a QA’s responsibility, but how do you grade those things?  Even worse, a manager of a large group of developers can’t be sure which tester caused which improvements or if it was indeed testers who did it.
  • Whittaker asserts automation needs to be backed up by manual testing, and I agree that automation has a great deal of limitations, and probably needs another 30 years (if ever) to mature before it really provides what has been sold to upper management.
  • The comments on correcting Microsoft’s quality problems were some of the better insights, yet the “how are we going to test this thing” with regard to design, doesn’t get  addressed often enough in the book.