Wednesday, October 29, 2014

In consideration of ISO 29119

I have been aware of this debate ever since attending CAST 2014, but I've not been quick to sign.  I wanted to investigate and see what other viewpoints there might be.  To quote Everett Hughes, a sociologist:
“In return for access to their extraordinary knowledge in matters of great human importance, society has granted them a mandate for social control in their fields of specialization, a high degree of autonomy in their practice, and a license to determine who shall assume the mantle of professional authority. But in the current climate of criticism, controversy, and dissatisfaction, the bargain is coming unstuck. When the professions’ claim to extraordinary knowledge is so much in question, why should we continue to grant them extraordinary rights and privileges?”
This is a serious question, one for which I appreciate that a standard might seem like a correct method for 'proof' of our professional authority.  However, even if the standard is in fact a professional guide, I don't think anyone but practitioners can judge its value.  As someone who has had a keen interest in trying to understand this standard, and not wanting to judge too quickly, I have tried to get a hold of a great deal of material about it.  I have engaged with some of those whom disagree with my point of view, such as professional tester.  In fact, they have published my response in there Oct. 2014 issue (albeit with some minor mistakes from my draft).  I looked into purchasing the standard.  I have looked for those who are pro-standard and what they've had to say about this movement as well as what they have said in the past.

That being said, I'm not convinced we actually do know what a standard should look like, much less if this is the standard we need.  Maybe it is what we need, but I strongly doubt it.  I think we are likely many years or decades away before we will even be able to claim we have true repeatable practices oriented towards different contexts, assuming it is possible.  I am unable to judge at present what the standard does say, as the standard and the standard body's work is less than transparent.  I've been forced to sign up with personal details in order to read documentation about the standard's creation, although that has been changed since (NOTE: The file name also changed, including a date, which makes it hard to know if anything has changed, as there was no change log as of 10/26/2014).  The standard requests I pay in a currency that I don't use.  I've signing the petition to withdraw the standard not because I know it is wrong, but because I can't tell, which makes it useless at best and dangerously wrong at worst.  What I can tell is that the standard's various author's other documents around the standard demonstrate what I consider to be confusing if not out right contradictory statements, making me doubt the actual standard.  That breaks the social bargain us professionals have that Everett Hughes so eloquently described.

Perhaps you could say that I should have been personally involved in the standard, and that is a valid complaint.  However, I'm not aware of anyone particularly reaching out to the AST, nor have I heard of it from any other group except for James Christie's talk in 2014.  I have attended both AST sponsored and non-AST sponsored conferences for years, so this isn't a case of willful ignorance.  This is my first chance to review the material and process, yet I have not found the process particularly open or transparent.  I see claims of no best practices and claims that the standard will create best practices.  I have found so many confusing statements by the standard's body that I must conclude that the standard should be withdrawn until it can be thoroughly reviewed and modified, if it can even be modified into something useful.

Even ignoring past statements, the recent defense of the standard creates questions.  One of the easiest and most obvious to consider is who wrote the standard.  Then there is the question of who pays for the creation of the standard?  Well clearly this was not just a labor of love, as Dr. Reid says the costs of development have to be passed on to the customer's of the standard.  I should note, my blog makes me no money and I am not a consultant so I have little incentive to make money speaking about this.  I made no money in writing my letter to the editor, and I certainly don't demand you pay to read my work.  I am not discounting the cost of writing the standard, just simply saying that if you plan on having expenses paid by publishing your work, you are not simply doing this out of the kindness of your heart.  There is lots of analysis that could be done just on the defense of the standards alone, but is outside the scope of this particular post.

One of the oddest and most compelling arguments for both sides is from Rex Black, in which he notes that about 98% of all testers won't care one way or another.  I think this is true, which makes the standards mostly not matter, but it also means that those 98% who are silent count on us to ensure these are the right standards, lest they become popular and that silent 98% ends up forced into using them.  I fear that this non-involvement is also further evidence that we testers as a group are not acting like a profession.  It isn't that we don't claim to have "extraordinary knowledge" and Dr. Reid at least seems to argue we mostly agree on this knowledge, but rather the majority of people don't feel any need to actively participate.  I realize this might be an argument for why we need a standard -- to show the disinterested the 'right' way to test, but to me it seems to indicate just how young our industry is and seems to me that shows why we aren't ready for a standard.

Even if the ISO body decided these documents ultimately should stand, the objections of the AST/Context Driven community need to be noted in such a standard.  Furthermore, making the document open will go a long way in allowing the community to discuss this document beyond the smaller standard's committee.  I recognize that ISO needs income to maintain itself and won't publish them for free for everyone, but certainly some sort of 'for individuals, not corporations' license could be used (and I don't mean the sort of non-sense Matt Heusser describes).  Finally, if this is an attempt to demonstrate our commitment to professional testing, then it needs to be accessible to our community.  The work needs to demonstrate it's value rather than being buried away inaccessible to those who would use it.

Tuesday, October 14, 2014

Noise to Signal: An Attempt to Understanding What Smart People Have To Say

Where I've been


Frankly, I don't know how to start this thing.  Perhaps it belongs in my bedroom when I was in high school writing little programs.  Perhaps it belongs just a few years back.  I have and am constantly looking to, in some way, do better for whatever "better" is for me.  I don't pretend to have answers for everyone or even that there is a static answer for myself.

In my bedroom I had finally found an interesting productive task.  I was learning to program.  I would search for hours for that one little magic bean to get me moving along again.  To use a game metaphor, I was leveling up in my knowledge at a rapid pace.

At my first testing job, I did firmware and hardware testing.  I sat around, watching machines work and sometimes break down.  As there was a large hardware component, I found quickly there was little to learn in rapid fashion, so instead I leveled up my political and social abilities, talking with the leads, learning how the organization worked.  I became a lead, and saw that it too was a trap, with just more paper pushing, so I found a different area to do more software-oriented testing.  I learned about LDAP and Kerberos and, and and ....  I was on the path to leveling up.  I was still programming in my bedroom, still writing games, and later different applications I found a need for, or an interest in developing.

Skipping a little, I started developing my own automation.  I built framework after framework, learning what things I didn't like and what I did.  I started in on teaching people.  I have mentored several people, sometimes more successful than others.  I leveled up in different areas around testing and development.  I had built my own testing tool, which admittedly has been neglected for years at this point, but I gained a lot of insight into the underbelly of an automation tool.  I even did market research to see if it was worth starting my own company.  I didn't start my own company for reasons not within scope of this blog, but it appears to have been a wise choice.

All of these things I did on the basis of my judgment and learning.  Isaac might have been a something of a mentor, I did read a few books, but mostly I did it on my own.  The last few years I have been looking at where to go from 'here'.  I have done most of the interesting difficult forms of automation, I have even worked towards pioneering my own methods and frameworks.  So in looking for what's next, I've been trying to learn from smart people and what they have to say, and not just within testing.

What do smart people say?


If you Google about what smart people know or do, you are bound to find pithy pieces of generic advice.  Not to say there aren't perhaps hidden gems in this, but these gems are minor.  Frankly in the world of wisdom, anything involving a list seems like someone who doesn't know what they are talking about but hope there work gets hits.  Add exclamation marks for further effect, and it feels less like real data and more like an advertisement.  Let me give you some examples:
  1. Work Hard!
  2. Don't Quit, Even When It's Hard!
  3. Never Stop Trying!
  4. Ignore Everyone Else and Do What You Think Is Right!
  5. Don't Look Back!
These are all easily disqualifable, be it too generic to be useful or falsifiable (E.G. Ignoring your past makes you a future fool ; another pithy statement, but sometimes true).  So seriously guys, how do I get better?

Well let me give you something that looks like a list but is rather a set of sources I have considered.  These are not quotes, but roughly the idea I took from them.
  • Don't set goals, use systems. - Scott Adams
  • Find things that don't scale in a linear fashion. - Scott Adams, Randy Pausch
  • Change is inevitable. - Robert Heinlein
  • Progress comes from people trying to solve a problem in a more effective/lazy way - Robert Heinlein, Paul Graham, Dr. Richard Hamming
  • Luck matters, but only a little - Too many to list
  • Knowledge and productivity compounds, the earlier or more you learn, the more you'll know later. - Dr. Richard Hamming, Steve Yegge
  • I have managed doubts in what I 'know', and that is okay. - Dr. Richard Hamming, James Christie
  • People's requests and errands take time, up to all of it. - Heinlein, Paul Graham
For those interested, I will include some links at the end of this blog for where I got some of my ideas.  I will stop there, because this isn't meant to be a final list of smart people's knowledge. Instead, I'm going to talk about the challenge Heinlein, Dr. Hamming, Paul Graham and Steve Yegge have all presented in some way or fashion.

Do you know what problems matter?  Why aren't you working on them?


After thinking about this mind-blowing set of questions, it makes me ponder what answer can I give?  I could give answers for my own company, but that feels too small.  I could give answers to the world's problems, but that is too big.  So maybe I can give answers around my industry, with a focus on testing:
  • Why do testers leave at year 5?  
  • Why do automation efforts fail so often (and is that related to people only staying in testing for 5 years or less)?  
  • How should we be teaching the next generation of testers?  Do we have accessible training designed to get people to the 5 year mark?  What about beyond that?
  • Why is only a small percentile of humanity appear useful to me (and am I in that percentile)?  
    • For that matter is my definition of useful a good and valid definition and would investigating that provide me with an interesting problem?  
  • How do we extract the most useful data in the world and provide it to the people who care?
  • How can I get machines to solve most basic problems?  
    • What do we do with those who were working before these problems were automated?
  • How do we effectively train the next generation so they don't repeat the same mistakes as the last generation?
  • When are particular methods of testing more or most productive?
  • How can we significantly improve testing without additional resources (E.G. Tools)?
  • How do we improve or prevent the conditions for sweatshop-like development companies?
  • How do we improve security in our software development?
This list was developed over one night and is not complete, but the important question to me is am I working on them?  Any of them?  However, I didn't write this blog post for me.  The real question is...  Are you working on problems that matter?  And if not, why not?  If you are working on important problems, then let's talk!



[Some] Sources:
http://www.paulgraham.com/hs.html
http://www.paulgraham.com/hamming.html
http://www.cs.virginia.edu/~robins/YouAndYourResearch.html
http://www.amazon.com/How-Fail-Almost-Everything-Still/dp/1591846919/
http://steve-yegge.blogspot.com/2008/06/done-and-gets-things-smart.html
http://www.youtube.com/watch?v=ji5_MqicxSo
Edit, Additional Link: http://www.youtube.com/watch?v=vKmQW_Nkfk8

Thursday, October 2, 2014

Societal Norms, Software Development and Culture

Preamble

I don't think that everything we do and consider in regards to test should actually revolve around testing.  That is to say, while testing is a primary consideration, it is just one of many considerations.  For example, making our team as optimal as possible might be a consideration.  Say that testing a particular feature would set back the sales group who are trying to do a demo of the product in the environment you are testing on.  Now not testing might in fact be considered an important activity in your testing career.  This post is simply about factors that while are very important in software development, they are not limited to that subject.

I personally tend to read many different and varied development related blogs, and you will find I write broad posts, posts that are not directly tied with testing, but certainly can be applied to testing.  This is one of those; you have been warned.  I have held this article for many months while I have struggled with the words, which I am still not completely happy with, so I hope you forgive the age of the linked article.  In my reading of the post Games Girls Onions, a post specifically about game development, I found some interesting questions I have asked in my career.  I think I can perhaps boil it down to two questions:

1. How do I deal with individuals?

2. How do I deal with groups of individuals?

This particular person found some of the interactions of her male co-workers to be sub-optimal, but the question is, did they intentional ignore those questions, or were they trying to find the right balance between the group and the individual?  Furthermore, if that balance is wrong, is that a cultural thing or a issue with a particular person?  Last but certainly not least, if some particular external group dislikes what the culture does, is it the individual or the group that needs to change?  Lets try to dive into these questions.

In my personal dealings, I have found that each and every individual is different.  That is not news worthy of course, but it is worth thinking about.  I have dealt with thousands of different people, and each one of them had some subtle variation to them.  However, most high level personality traits could be categorized, and those categories can play a part in how you deal with someone in a new situation.

Complexity


Consider this example, you know your over-zealous co-worker who cares a great deal about usability.  Even though you don't know for sure, you can guess that like other testers you've dealt with in the past, they too will enjoy working in an area they are passionate about.  Offering to trade your UI piece for their API piece might be a good idea if you don't have any particular passion for UI testing.  This of course assumes your boss is okay with it, which makes this at least a small group choice.  Since you know the last few times you have juggled work around, she was totally good with it, so you don't even bother to ask first and just do it.  Lets pause here and consider.  You choose to do the thing that made your co-worker happy while not bothering your boss, as they are likely busy.  From an individual basis, you made your co worker happy.

You didn't consider the team, as you don't 'report' to the team.  Now in a culture that appreciates employee happiness and efficiency, this might work out fine.  In a culture that rather have consistency and order, this choice may be a problem, even when no one is directly affected.  Your boss might respect your choice of taking initiative, but still might chew you out because the group or culture doesn't embrace that style of work.  Maybe the guy you swapped with is disliked by one of the developers on the UI team (as he is too passionate), so this makes things worse for the team.  But who decides these things?  Human activities are really complicated and to put blame on any individual is likely unfair.

Norms


Social norms start to show up in a particular culture to define what is or isn't acceptable so that way a group knows how to react and what is acceptable.  I realize saying this might offend someone, but when you try to break social norms, expect the society to attempt to break you.  In the culture of development, there is an assumption that it was a male dominated activity.  Like it or not, that is presently, in most countries, a social norm and those who break that norm are more likely to feel social pressure.  Is it natural?  Norms are natural in human society, this particular norm is just cultural.  I recently heard in a BBC radio interview that one country (no citation, as I can't find the details) has about a 1:1 ratio of women to men in software development.  To consider a different culture and norm, Victorians expected a particular style of dress to fit into a particular class.  If you didn't fit into that dress, you failed their normative test and you were looked down upon.  Now that norm is broken and so we can now wear Jeans acceptably.  On the other hand, in my culture, I can't wear slippers into work everyday without being looked at strangely (if not fired, depending on which job I look back at).

Back to considering the article, the developer talks about her nice window office -- a thing that is usually reserved for higher level developers.  Maybe that isn't the case at her office, but to me that is automatically a status symbol.  Now she notes that one time, a weird incident happen where a drunk co-worker showed up and talked to her in a way she didn't like for 20 minutes.  She talks about being treated differently, because she is being classified in what is defined as a social norm that she did not agree to.  The problem is, we as a society don't have sophisticated categories in our culture which model real human beings.  It is difficult to know exactly how to handle a female developer who accepts crude jokes, wants to be looked upon as socially confident, but then feels threatened by being alone with a drunken coworker.  It is not that I don't appreciate the dichotomy, how situations can be uncomfortable, nor how the details might be important in giving me a more solid opinion.  In culture, norms tend to be broad heuristics and in order to change the cultural and norms, it takes a long time and requires us to make mistakes along the way.

Being a Fraud


In fact, right now, we have a very cold, unfeeling society.  Consider Kevin Fishburne's comment in that post,
From the mortgage industry to pizza joints it's been my experience that regardless of how good or bad your relationships appear at work once you walk out that door they're over. I liken being at work to acting in a play; you're not being you, but what you think you need to be at the time.
 I'm sure there are plenty of exceptions where great lifelong friendships are formed on the job but it hasn't happened to me, sometimes much to my surprise. Not sure if this is just me or an actual phenomenon, but if the latter it may have some bearing on why people act differently at work then they would elsewhere....
The idea that we all follow social norms, faking our way through the work experience appears to not be particularly a-typical.  The fact that the author feels the need to complain about these issues shows she is fighting at least one social norm.  As a tester, it is not unreasonable for us too to want to test our current social situations and attempt to improve them.  One item that I agree with the author on, but am not sure is a 'bug' is the feeling that you are a fraud.  To take a few quotes from the movie Art & Copy,
The frightening and most difficult thing about being what somebody calls a creative person is that you have absolutely no idea where any of your thoughts come from, really. And especially, you don't have any idea about where they're going to come from tomorrow. - Hal Riney

I think most of the creative people are so damn insecure that they want to think they know everything, but they know deep in their hearts they're just in deep trouble from the minute they get up in the morning. So if you can tell them "that's what you're supposed to be", that's kind of liberating. - Dan Wieden

When we are often not even genuine with ourselves, how can we be honest with others?  Dunbar's number (and others) suggests we can't care about more than about 300 people, with the number more likely 150 people.  This isn't just other people, this is you too.  This isn't just you, this is other people as well.  Keep that in mind.  Since our capacity to care about humanity is so low, our ability to repair or change an entire culture is extremely limited.  Thus our social norms are not and can't be designed by any particular person for the masses.  On the other hand, I can only keep track of so many people's personal preferences.  When a culture says something like, "Dirty jokes are okay at work" and the broader culture says "Dirty jokes are never okay when a woman is around," what do you do?  Perhaps your mind checks to see if you know the woman and decide based upon that?  Perhaps you test them with a joke?  What happens if that one joke gets you called in front of HR? I recall one man being written up for calling a group of testers 'test monkeys' (he too was a 'test monkey' in his own opinion), but that insulted one man's sensibilities.  You can't please everyone.

We Are All Unique


The alternative is to say we are all unique.  This idea of course runs counter to Dunbar's number that I can't handle everyone and have to limit myself to a certain set of people whom I use categories for.  Even if we are all unique I can't know how to start with each unique personality.  I start to develop rules, like "this group seems to often act like X" and apply that, but it takes time.  I can't start fresh, not knowing anything.

For example: Does silence means yes or no?

If you answered the question you are wrong.  If you remained silent you might too have been wrong (unless silence means I don't know).  As an example of yes, consider the political process in the EU.  Wiki itself notes that silence doesn't always mean consent.  In one of my Psychology text books back in college, I recall read that in the middle east, a pilot attempted an unscheduled landing and asked for permission to land.  He got no response back and assumed that meant yes.  He got into some rather big trouble when he tried to land and they shot at him, because silence meant no.

If you attempt no assumptions ever, you have to relearn everything every time, and with reading other's writing, you can't even know if the person you are learning is the same each time, as a persona does not equal a person (as Ben Franklin demonstrated).

At least in one example in the article, the author mentions how swearing is fine by her.  One of the commenter noted that she did not feel the same way.  If you can have people upset both ways, you bet there is a third way.  Hell, even asking, which feels a little odd, might upset someone who assumes that the matter is already culturally settled.  Note how I wrote hell just now.  I wonder if I upset any readers or caused them to pause.

Impossibility to Adapt


If you can't adapt because everyone's situation is too dynamic and everything depends, what do you do?  Since we can't adapt and culture is impossible to change by one person, we must suffer, with those who fight to change the situation slowly, making sacrifices for what they believe in.  I don't mean to sound harsh, but the only ways to change a culture is to work really hard at it and often suffer slings and arrows.  Sometimes waiting for others to move on or die off helps, but that takes years.  Alternatively, one can leave or build their own culture.  Basically, no matter which option you take, expect that it will take a lot of work!

Thursday, August 21, 2014

Words of the Week: Critique and Criticism

Preamble / Ramble


This is a first for me and word of the week, as I typically consider a word without much personal context, but this week I am going to include a context. While at CAST this last week, I heard some valuable input regarding my blog. No matter how I get feedback, I try to take it and make it useful. I tend to be direct with my feedback, but for some, that can sound harsh, so this week I’m going to consider two different words. If you don’t want to hear the backstory, just skip down to the next section. However, for those that care, I do want to give some context and clarity in my own words. I have now heard from multiple sources that they felt my writing reflected a negative viewpoint towards the Software Testing World Cup 2014. I think that to be an inaccurate statement. I did have some issues with it, and I felt I was doing a deep dive analysis of both the pros and the cons of the contest, up to the point I had gotten. In assuming that I was writing a report about my experience with the contest, going into both the good and the bad. I also wanted to talk about what I could gain from the contest and what could be done better. In my second part I wanted to consider what the judges had to say, so I could learn where to improve as well as note what additional feedback would be nice to have. I understand that this is a contest, but I think the more important piece is what can be learned from it. To be fair, perhaps the organizers didn’t want what I had to say, as they actually asked for:

“At the moment we are very curious about your personal STWC experience. We would love to read blog posts or any other kind of written reports about your participation. It doesn’t matter if you write from your personal point of view or from the perspective of your whole team. We are interested in hearing how much fun you had during the competition, what you have learned from this experimental challenge, and what the difference is between testing while having fun together with friends and testing at work. We are also interested in how you communicated with your team members and what your team’s strategy was. Did you start testing right away or did your team follow a strategy? How did you create that strategy? What was your biggest challenge within your team, what kind of bugs were you able to detect and how did you like working with the Agile Manager tool?”

I took this to mean they were excited to receive any feedback regarding the design of their contest. They also wanted to hear what we did and why we did it. I heard that they were open for a critique of the process so that they could attempt to improve it. Furthermore, we were not told not to post scores, as we felt that it was important that those who participated could compare scores and learn where they could improve.  In fact I encouraged others to post scores, but no one actually did so. We also thought it might add value to future judging to see a relatively unemotional reaction to the scores given. We made the judge’s scores more anonymous as it felt important not to provide any identifiable information. However, it seems some saw what I thought was a critique others felt was criticism. This leads me to the two words, critique and criticism. Having re-read my content I can see that it could be read as criticism, and while I won’t take back my words, I will keep that in mind for the future. So on to the words of the week.

Back to our Regularly Scheduled Program


First allow me to define these two words in the way I think of them, without doing much research nor using Google to define the words. I think of criticism and critique as similar, but with a difference in tone and intent. To criticize someone is to give back negative statements, with either the intent to hurt/harm, with no positive points and to lack any possible ways to improve, be they implicit or explicit. Calling someone stupid is to be critical of someone’s intellectual factuality. It is likely intended to harm, has no positive points and does not hint at a way of improving. To critique is to provide feedback with the intent to improve a person or person’s work in some way. There is a mix of some positive points in the comments as well as negative comments. Gushing praise or all negative points is probably not a critique. Furthermore, at least some of the statements need to be actionable. Critique also feels very student/teacher-y. Criticism on the other hand seems like it refers to bullies and angry people shouting at each other. I want to capture this in a readable way for later usage:

Criticism Critique Gushing Praise

  • Negative comments
  • Intended for harm/hurt
  • No method for improvement provided.
  • Negative and positive comment mix
  • Intended to improve the content, person(s) or future
  • Provides methods for improvement.
  • Positive comments
  • Maybe intended for manipulation (even if meant to be ‘improve another person’s day’)


Now let’s see how I did. There are a lot of different people’s views on these words, so let me try to capture just a few of them.

Various versions of Critique and Criticism


According to the presently [August 20th, 2014] most popularly rated stack exchange view the difference is zero. They point out how in 1960’s academics started to try to use critique as a form of analysis rather than meant to censure, but that didn't stick.  They cite several modern examples of mixed usage of the terms. In comparison, Paul Brians, the author of Common Errors in English Usage, says:
“Josh critiqued my backhand” means Josh evaluated your tennis technique but not necessarily that he found it lacking. “Josh criticized my backhand” means that he had a low opinion of it.
Clearly their is some difference of opinion on the meaning of the word. In Writing Alone, Writing Together; A Guide for Writers and Writing Groups by Judy Reeves, there is a interesting comparison of criticism and critique:

“[1]Criticism finds fault/Critique looks at structure
[2]Criticism looks for what's lacking/Critique finds what's working
[3]Criticism condemns what it doesn't understand/Critique asks for clarification
[4]Criticism is spoken with a cruel wit and sarcastic tongue/Critique's voice is kind, honest, and objective
[5]Criticism is negative/Critique is positive (even about what isn't working)
[6]Criticism is vague and general/Critique is concrete and specific
[7]Criticism has no sense of humor/Critique insists on laughter, too
[8]Criticism looks for flaws in the writer as well as the writing/Critique addresses only what is on the page”

(NOTE: I numbered the following quote for later convenience)

It is not completely clear some of the differences between criticism and critique, as I can find fault in the structure of a work, which leaves me in no mans land. Wait… wait…

I want to analyze my last sentence, using the rules listed above. In the second rule, I clearly have a criticism of the writing. The the third rule says if I had asked in the form of a question, if I had said “It appears not completely clear some of the differences…” and perhaps ended with a question mark it would have been a critique rather than a criticism. That appears to me to mean the syntax matters more than the semantic content in the author’s opinion?* Oh, this gets us to rule four. While I made my statement objective and honest, it might not be considered kind. I can’t tell if my statement could be seen as negative or positive. It certainly isn’t positive, so I’m going to read between the lines and call it criticism. I think my statement seems concrete and specific, even if no exact example was given. However, it isn’t full of humor, nor is most of my writing.** The final dimension is about the author. With no reference to the author, I think I am addressing only what is on the page. Ultimately it feels that this set of tools has limited value as it takes too much analysis and when you have 5 / 8 on the criticism side, is it criticism? To the definition's benefit, it does try to capture the concepts I stated early about providing some positive points and providing negative points without beating the subject up. It adds to my definition attempt that you should be addressing the content rather than external influences and adds the idea that the critical comments be concrete. However, it fails to address the possibility of simply gushing rather than capturing the areas of improvement.

Obviously my example is contrived, but look at the above paragraph as whole (yes, I’m getting meta) I would call it a criticism using the rules Judy Reeves developed, yet it wasn’t intended to be. So it feels wrong in some way or perhaps I’m simply wrong. In my mind, she is looking for more positive views rather then stressing the analysis. That is to say the message is less important than the presentation. In my opinion, I’m critiquing the content, because I follow my own rules as well as the new rules I have developed.


* Yes, that was intentionally sarcastic. I hope you can appreciate the joke.
** Ignoring the above *’d item. It wasn’t that funny anyway.*

To Sum Things Up


Some people don’t see a difference between criticism and critique. In more academic circles, it appears that a divide is seen, even if the divide is a little fuzz. Some see it as a difference of purpose, some see it as a difference of the message and some see it as a difference of the content.  How do I see it?  Well, let me capture the attributes I think matter:


Criticism Critique

  • Negative comments
  • Intended for harm/hurt
  • No method for improvement provided
  • May focus more on the authors over the content
  • "Censure, attack, abuse and name call"
  • Negative and positive comment mix
  • Intended to improve the content, person(s) or future
  • Provides methods for improvement
  • Address the content rather than the authors
  • Concrete examples
  • "Deep analysis of other people's work"

In reviewing my own work, I think it falls under the label critique, but I can see how others might read something into it.  Written critiques are hard to convey personality.  You can see that with my footnote joke above.  In going so meta, I was intending on injecting some levity into an otherwise highly intellectual process.  But others might see it as 'unprofessional' or they might read my words as snarky.  In getting my article reviewed, Isaac said he thought it was a little on the snarky side, but that in his view it was just right to make the point.  Outside of getting my content reviewed, I am not sure I can defend against hurting people with a wide audience and perhaps that is Judy's point.  By being kind, you avoid the question completely.  Do you have a critique or feedback for me?  Feel free to post in the comments.

As a sort of post script, I happened upon this site after I wrote this article.  As I didn't want to shoe horn it in, yet felt it was of some value, so I put it here.

Friday, June 27, 2014

My Current Test Framework: Testing large datasets

I recently wrote how I felt few people talk about their framework design in any detail.  I feel this is a shame, and should be corrected as soon as possible.  Unfortunately, most companies don't allow software, including testing automation to be released into the public.  So most of the code we see is from consultants with companies occasionally okaying something.  In my case, I did something like a clean room implementation of my code.  It is much simplified, does not demonstrate anything I did for my company nor does it directly reference them.  It is open source and free to use.  Without further delay, here is the link: https://github.com/jc-d/Demos.

It is in Java and was intended as part of a demo for a 2 hour presentation.  Because of the complexity of the system, I'll write some notes about it here.  I used IntelliJ to develop this code and recommend using it to view the code.  It does use Maven for the libraries, including TestNG which you can run from IntelliJ.  Many of the concepts could be translated to C# with little difficulty.

So what does it do?  It demonstrates a few different, simple examples of reflections and then a build up of methods for generating test data in reflective and possibly smarter fashion (depending on context).  I'm sure your sick of hearing about reflections from me, so I'll try to make this my last talk on them for a while, unless I come up with something new and clever.

As a brief aside, Isaac claims that while this is a valiant attempt to write out a code walkthrough, it really needs to be a video or audio recording.  Perhaps so, but I don't want to devote the time unless people want it or will find it useful.  As I have my doubts, I'm going to let the text stand and see if I get anyone requesting a video.  If I do maybe I'll put some time into it.  Maybe. :)

Now on to the code...

First the simple examples which I do use in my framework but not as simple as this.  The two simple examples are DebugData and ExampleOfList and both are under test/java/SimpleReflections.  DebugData shows how you can use reflections to print out a simple object one level. down.  It 'toStrings' each field in the object given.  Obviously if you wanted sub-fields that would take more complex code, but this is often useful.  ExampleOfList takes a list of string and runs a method on each item on the list and returns back the modified list.  Obviously this could be any command, but for simplicity of the demo I limited it to methods that did not take arguments.

Now all the rest of the code is around different methods for generating data.  I will briefly describe each of them and if you want to you can review the code. 

The HardcodedNaiveApproach is where all of the values are hard coded by using quoted strings.  E.G. x.setValue("Hard coded");  This is a good method for 1-3 tests but if you need more you probably don't want to copy and paste that data.  It is hard to maintain, so you might go to the HardcodedSmarterApproach.  This method uses functions to return objects with static data so you can follow the DRY principle.  However, all the data is the same each time.  So you add some random value, maybe append it to the end.  The problem is what are your equivalent class values?  For example, do you want to generate all Unicode characters?  What about the error and 'null' characters?  Are negative numbers equally valid to positive numbers.  If not, then your methods are less DRY than you might want as you will need different methods for each boundary, if that matters.  Also you are writing setters for each value, which might fail when a new property is added.  We haven't even talked about validation yet, which would require custom validators based upon the success/failure criteria generated by the functions you write.  That is to say if you generate a negative number and it should fail for that, not only does your generator have to handle that but your validator does as well.   What to do?

Perhaps reflections could help solve these problems?  The ReflectiveNaiveApproach instead uses the typing system to determine what to generate for each field in a given class.  An integer would generate a random integer and a string would generate a random string.  We know the field name and class type so we could add if statements for each field/type but that puts in the same maintenance of new properties we had with the hard coded approaches.  If we didn't do that we could still handle new properties assuming we knew how to set the type, but it might not fit the rules of the business logic and we have no way to know if it should work or not.  For fuzz testing this is alright, but not functional testing.  Is there any solutions?  Maybe.

The final answer I currently have is the ReflectiveSmarterApproach.  In effect when you need to generate lots of different data for lots of different fields, you need to have custom generators per class of fields.  What is needed is an annotation for each field telling it what needs to be generated.  An example of that can be found in the Address class.  Here is a partial example of this:

public class Address {
 @FieldData(dataGenerators = AverageSizedStringGenerator.class)
 private String name;
 @FieldData(dataGenerators = AddressGenerator.class)
 private String address1;
 //...
}

Now let us look at an example generator:


public class AddressGenerator extends GenericGenerator {
 @Override
 public List<dynamicdata> generateFields() {
  List<dynamicdata> fields = new ArrayList<dynamicdata>();
  fields.add(new DynamicData(RandomString.randomAddress1(), "Address", DynamicDataMetaData.PositiveTest));
  fields.add(new DynamicData("", "Empty",
   new DynamicDataMetaData[] {DynamicDataMetaData.NegativeTest, DynamicDataMetaData.EmptyValue}).
   setErrorClass(InvalidDataError.class));

  return fields;
 }
}


This generator generates a random address as well as an empty address.  It is clear one of these addresses is valid while the empty address appears to be a negative test.

Through the power of reflections you can do something like this:


List<DynamicDataMetaData> exclude = new ArrayList<DynamicDataMetaData>();
exclude.add(DynamicDataMetaData.NegativeTest);
ReflectiveData<Address> shippingAddress = new CreateInstanceOfData<Address>().setObject(new Address(), exclude);


The exclude piece is where you might filter out generating certain values.  Say you want to only do positive (as in expected to be successful) tests, you might filter out the negative tests (those that expect not to complete the task and possibly cause an error).  The third line generates you an object with all the properties that have the attached annotation and values.  Now this does not handle new fields automatically but it could certainly be designed to error out if it found any un-annotated fields (it is not at present designed to do this) and if you embed the code in your production code, it would be more obvious to the developer they need to add a generator. 

Now how do we pick which value to test when we could test with the empty address or a real address?  At present it picks it randomly because according to James Bach, random only takes roughly 2x to get equal coverage to pair wise testing.  Since we know all about the reason for generating a particular value (what error it would cause, etc.) we can at run time say how the validation should occur.  The example validation is probably a bit complex, but I was running out of time and got a bit slap-dash on that part.  One issue with this method is it is hard to know what your coverage is.  You can serialize the objects for later examination and even create statistical models around what you generated if needed.

Summary

Obviously this is a somewhat heavy framework for generating say 20 test data values.  But when you have a much larger search space that approaches infinite this is a really valuable tool.  I have generated as many as 50 properties/fields about 400,000 times in a 24 hour period.  That is to say, generating roughly 400,000 tests.  I have found bugs that even with our generator would only be seen 1 : 40,000 times and would probably have never been found in manual testing (but would likely be seen in production).  The version I use at work has more than a years worth of development and research, supporting a lot more complexity than exists in this example, but I also don't think it could be easily be adapted as it was built around our particular problems.

This simple version can be made to support other environments with relatively little code modification.  It took much of the research and ideas I had and implemented in a simpler fashion which is more flexible.  You should easily be able to hook up your own class, create annotations and generates and have tests being generated within a day (once you understand how).  On the other hand it might take a little longer to figure out how to do the validation as that can be tricky.

One problem I have with what I have generated is there is no word or phrase to describe it.  In some sense it is designed to create exploratory data.  In another sense it is a little like model driven testing in that it generates data, has an understanding of what state it should go to and a method to validate it went to the correct state.  However it doesn't traverse multiple states and isn't designed like a traditional MDT system.   Data Driven Testing describes a method for testing using static data from a source like a csv or database.  While similar, this creates dynamic tests that no tester may have imagined.  Like combinatorics, this creates combinations of values, but unlike pairwise testing, the goal isn't just generating the combinations (which can be impossible/unpractical to enumerate) but to generate almost innumerable values and pick a few to test with, while enforcing organization of your test data.  This method also encourages usage of ideas like random values while combinatorics is designed to have a more static set of values.  Yes you can make combinatorial ideas works with non-static sets, but it requires more abstraction (E.G. Create a combination of Alpha, Alpha-numeric, ... and this set of Payment method, now use the string type to choose what generate you use) and complexity.  Finally combinatoric methods can have difficulties when you have too many variables, depending on implementation.  This is a strange hybrid of multiple different techniques.  I suppose that means it is up to me to try to name it. Let's call it:  JCD's awesome code.  Reflective Test Data Model Generation

I would say don't expect any major changes/additions to the design unless I start hearing people using it and needing support.  That being said I love feedback, both positive and negative.

While researching for this article I came across this which is cool, but I found no good place to cite it.  So here is a random freebie: http://en.wikipedia.org/wiki/Curse_of_dimensionality

Thursday, June 19, 2014

What is the Highest Level of Skill in Automation?

Thanks to Robert Sabourin for generating this topic.  Rob asked me roughly, 'What in your opinion is the highest level of skill in automation?'  He asked this to me in the airport after WHOSE had ended, while we waited for our planes.  It gave me pause in considering the skills I have learned and help generate this post.

Let me make clear a few possible issues and assumptions regarding what the highest level of skill is in automation.  First of all, I think that there is an assumption of pure hierarchy, which may not exist.  That is to say, there might not be a 'top' skill at all or the top skill might vary by context.  So I really am mostly speaking from a personal level and with my own personal set of automation problems I have faced.  When I answered Rob's question in person, I neglected to add that stipulation.  The other possible concern is that the answer I give is overloaded, and so I will have to work on describing the details after I give the short answer.  Without making you wait, here is my rough answer: Reflections.

What are reflections?

In speaking of reflections, you might assume I am speaking of the technology, and for good reason.  I have spoken on them many times in this blog.  However, that is just a technical trick, albeit a useful one. I am not talking about that trick, even if the comp-science term 'reflections' is part of the answer.  In speaking of reflections, I mean something much broader.

There is the famous "thinker" sitting on his rock just pondering is much closer to what I had in mind.  But you might say, "Wait, isn't that human thinking?  Isn't that critical thinking or introspection?"  Yes, yes it is.  What I mean by reflections is the art form of making a computer think.  While a computer's intelligence is not exactly human intelligence, the closer we approach that vast gulf, the closer we are to generating better automation.

Most people might start to argue that requires someone with in depth knowledge or artificial intelligence or at least a degree in computer science or someone with a development oriented background.  Perhaps that is the logical conclusion we will ultimately see in the automation field, but I don't think that either an in depth knowledge of development or AI is required for now.  I know that you need to go to that level to start understanding this concept.

Instead, I think you need to start thinking of the automation in the way you think about writing tests.  In some ways this relates to test design.  Why can't the automation ask what am I missing? Why can't my automation tell me what the most likely reason a failure occurred*?  Why can't the automation work around failures*?  Or at the very least, ignore some failures so it isn't blocked by the first issue it runs into*?

* I've done some work around these, so don't say they are impossible.

Now that I have walked around the definition, let me define the reflections in context of this article.

Reflections:  Developing new ideas based upon what is already know.

An example

A good example for the need for reflections is the brilliant talk given by Vishal Chowdhary, in which he notes that in translations (and searches, etc), you can't know what the correct answer is.  You have no Oracle to determine if the results are correct.  Many words could be chosen for a translation and it is hard to predict which ones are the 'best'.  Since computer language translations are adaptive, you can't just write "Assert.Equals(translation, expectedWord)" with hardcoded values.  Since these values are dynamic, the best you can do is to use a "degree of closeness".  You see, they couldn't predict how the translation service would work because it has dynamic data and the world changes quickly, including new words, proper titles, et cetera.

So how do you test with this?  Well you can look at the rate of change between translations.  You can translate a sentence, translate it back and record how close it was to the original sentence.  Now track how close it is over time, with different code and data changes.  You could take translation string lengths and see how they vary over time and note when large deviations occur. There are lots of methods to validate a translation, but most of them require the code to reflect on past results, known sentences and the likes.  The automation 'thinks' about its past, and on that basis judges the current results.

Not to say some automation shouldn't be reflective.  For example you could hard code a sentence with "Bill Clinton" in it and check to make sure that it didn't in fact translate his name.  You could translate a number and check to see it didn't change the value.  You might translate a web page and check something not related to the translation such as layout.

Not just the code

In reading my blog you might assume that because I specialize in automation I think reflections is a code-oriented activity.  I do think that, but I think it applies more broadly.  When I write a test, I should be reflecting on that activity.  That is to say I should be thinking "Is that really the best design?", "Should I be copying and pasting?", "Should I really be automating this?", etc.  In always having part of my brain reflecting on the code, I too am write better code.  Hopefully between my writing better code and my code trying to do better testing using reflections, we do better testing overall.  This also applies to testing in general, with considerations around things like "That doesn't look like the rest of the UI." or "I don't recall that button there in the last build."

I have only scratched the surface of this topic and made it more specifically apply to automation/testing, but I think this applies to life too.  For a more broad look at this topic I would highly recommend Steve Yegge's blog post Gödel-Escher-Blog.  It will make you smarter.  Then next time you go do some automation, reflect upon these ideas. :)  And if you are feeling really adventurous, please put a comment about your reflections on this article here.

Tuesday, May 20, 2014

Software Test World Cup 2014 - Part 2

I wrote recently about the Software Test World Cup. In the last post I said I would post our score card and an analysis if one was possible. First I will give you the raw data, with some "Avg" columns removed. I have not intentionally edited the content other than making the judges more anonymous, format shifting from excel and removing extra columns involving averages and totals. Oh and I marked spelling mistakes with sic, which I just learned should be written with brackets, not parenthesis.

JudgesImportance of Bugs FiledQuality of Bug ReportsNon-Functional Bugs filedWriting/ Quality of Test ReportAccuracy of Test ReportBONUS: Teamwork/Judge interaction 0-10NOTES
A1616914141Bonus: Test Report/Bugs made me want to engage the customer. Several Usability issues.
B15151113131I found that the report did not flow well. I know many teams are expected to give ship/no ship decisions, this [sic] iritates me, i promise not to let it affect my [sic] juding.
C12111014121I like the test report. It gives practical examples, what could have been [sic] testeed for different aspects, e.g. Load, but the ship decision and the "major" bugs seem not to fit imo. bug spread is okay, they tried to consider disability issues (red/green vs. [sic] colorbling ID326)

First of all, I thank the judges for making not only numeric judgments, which is super hard, but also that they spent the time to write some comments. I appreciate that. However, in looking at the judge's comments it seems confusing. For example, C says they liked the report, but their score was no higher than the other reports. They complain that the major issues seem not to fit, such as crashes and the inability to order the product. Granted I don't have the full list of bugs we found, but I wonder what priority they were looking for. That is left to the reader's imagination.

Judges A and B were kinder score wise. Judge A gives us a bonus, but not the bonus we were promised by Matt Heusser, the product owner for answering a question in the youtube channel. Certainly the bonuses were not used how I imagined. I thought our conversations with Matt and the Product Owner would be part of that bonus, but it appears the judges ignored this. I did find it interesting that Matt said that usability would be part of non-functional [no citation, I'm not rewatching 3 hours of video]. We also asked for permission to do some load testing on the Snagit site but never got a response from the product owner, so we chose not to because of legal and ethical implications. It seems like that was not applied in the scoring, but maybe I am wrong. With Snagit as the tool to test there isn't a lot of non-functional testing to do. Judge B on the other hand was harsh in comments but gave good scores (relatively). I agree with Judge B that providing a ship/no ship is annoying, but that is what a conversation is for. We didn't get to have one of those, so we did the best we could.

It is interesting how diverse the opinions of the judges were and also how middle in the road most of their scores were (all items but the bonus were out of 20). Considering how highly we placed, I am guessing either the judges were never impressed or they found out how bad things got and a 10 / 20 really is more like a 15 / 20 relatively speaking. Finally, I promised our test report. Sadly I can't upload the thing to blogspot, so instead I am going to post it below, with some attempts to deal with formatting.


Functional Test Report

By: JCD, Isaac Howard, Wayne Earl, and KRB


Status:

Do not ship

Major Issues:

Undo does not always work, undoing the wrong thing. We found a good number of bugs, including multiple crashes; some seem more realistic seen than others. Bug 241 showed that webcam capture crashed on one particular Mac. When no webcam exists and you attempt to take a camera capture Snagit closes. The Order Now, Tutorial, Get More Stamps and New Output buttons on went to a 404 page. With Preferences open in windows 7, 8, the application refuses to take screenshots. There are about 20 priority 1-3 bugs, which suggests the application isn’t finished.

What Did Work:

We did a performance test with a small Win 8 with 4 gigs of ram and 1.8ghz cpu. It succeeded to take video at a viewable quality. The editor in most cases worked well. Basic usage in most cases works well. The mobile integration worked.

Misc:

We earned bonus points per Matt for giving advice on how to edit video (Jeremy Cd). We asked multiple times in the youtube channel if we could load test the system’s web site but could not get back a response from the judges/Matt. We chose not to ‘hack’ the system due to legal and ethical issues.

Limits of Testing:

- No Automation was generated
- State of Unit Testing is unknown (Customer’s don’t know about unit tests)
- We did even come close to hitting all the menu items.
    o We only have a rough set of tests of windows. Most of our testing was on Macs.
- We didn’t have the technical expertise in the system to capture logs.
- We didn’t have the time to capture the before and after
- We have limited experience with the SUT.
- We only tested the configurations provided. Other configurations of the system were ignored.
- Could not have a conversation with product owner on critical bugs after he left.
- Driver/Hardware testing is limited. We mostly have Maverick OSX.


How Testing Was Planned:

-Website:

As we have been informed that this is a highly hardware dependent product (video and screen capture), which is a commodity, we decided that actually looking at the website of the product might be as important as the product itself. Particularly since a user cannot determine the differences in quality between products without testing the products, the website becomes a primary concern. Particularly since we don’t know they support the types of computers we have, we might be forced to test the site. Finally, if the product has a price, making sure you can’t break the security and get to the download system for free is important.

-Load:

In discussing load testing, we have considered trying to record multiple youtube videos all playing at once, which will stress the hardware and will make for a great deal of variation of the recorder to capture. It might also add a lot of audio channels to record if required. Also using a TV-static screen will push the compression algorithm, as it cannot be compressed well, so we might also test with that. Finally attempt to use a slower system, with little ram and a slow hard disk with lots of CPU and disk usage to see if the recording fails.

-Mobile:

We will attempt to get our relatively few mobile devices to load the system if possible and do some basic usage.

-Usability:

If it is not easy to do the basics of screen/video capture, then users may ask for their money back or go find a different product.

-Feature comparison:
http://en.wikipedia.org/wiki/Comparison_of_screencasting_software#Comparison_by_features
http://lifehacker.com/5839047/five-best-screencasting-or-screen-recording-tools

-Functionality:
• Save a file with a good and a bad file name.
• Record and Capture on app basis, full screen, area
• Two monitors vs One Monitor
• Editing if supported
• Sound if supported
• Pan and zoom if supported
• Arrows, Text, Captions, etc.
• Transitions
• Competitive Intel:
    o http://www.techsmith.com/tutorial-camtasia-8.html
    o http://www.techsmith.com/jing-features.html
    o http://www.telestream.net/screenflow/features.htm
• Long time record (If possible)
• Upload Tools
• Formats supported
• Merging/Dividing recordings
• Tagging / describing the videos other than just the file name
• How easy is it to take the output and use it with another system (e.g. I want to quickly use screen shots and videos from this tool and add them to a bug I created in JIRA)
• Can you add your own voice to the recording? Like narrate what you are doing or what you expect via the microphone on the device you are using
• Can you turn this ability off so you don't hear Isaac swearing at the system or can you remove the swearing track after the fact?

Questions:
- Who are the stakeholders for this testing?
- What are the requirements? What is the minimum viable feature set?
- What is the goal of this testing (Important bugs, release decision, support costs, lawsuits, etc.)?
- Can you please give 3 major user scenarios?
- What is the typical user like? Advanced? Beginner?
- What are the top 3 things the stakeholders care about, such as usability and security.
- Can you give an example of the #1 competitor?
- What sorts of problems do you expect to see?
- Does the SUT require multi-system support? What systems need support (E.G. Mobile, PC, Mac, etc.)? Are there any special features limited to certain browsers/OSes? How many version back are supported?
- What is the purpose of the product? Do you have a vision or mission statement?
- Can you describe the performance profile?
- Is there any documentation that we should review?
- What languages need to be supported?
- Does the application call home (external servers)? Are debug logs available to us? Even if not, is anything sensitive recorded there? Are they stored in a secure location, either locally or externally?
- Do we know what sorts of networks our typical user will use to download the app with? Do we know what the average patch size is?
- Are there other possible configurations that might need to be tested?
- How does the product make money? What is the business case rather than the customer case?
- Should we/can we do any white box testing?
- Can we get access to a developer to ask questions regarding the internals of the system, code coverage, etc.?







Sunday, May 18, 2014

Software Test World Cup 2014 - Part 1

[Part 2 is now posted which includes the scores we received and an analysis.]

I recently participated in the Software Testing World Cup 2014. Just to get it out of the way, our group did place in the semi finals, but did not win. However, that is of less importance to me. Wait you say, what is the point of a contest if not to win. Well for me, I was more interested in getting feedback, learning from the experience and seeing what the idea was. Since this is our blog, not my resume, I'd rather talk about what we did than sell myself. What I am going to try to capture is the experience and outcome of the event.

In starting about the event, I should be clear, a large part of the event happened before the event even started. We were asked to prepare to: Interact with the customer, test the software, write bugs and come out with a report on what was found. We also were told the judging would be on: Being on mission, quality of the bugs, quality of the test report, accuracy, non-functional testing and interaction with the judges.

Knowing that, I did a few things. The first was I assumed it was likely to be web based, as most apps don't work on all platforms, so I started to enumerate what sorts of non-functional tests we could do. This turned out to be wrong, but it is better to be over prepared than not in my view. Then I looked into how to write such as report as I don't do those sorts of consultant-oriented paperwork. I have conversations with the developers and work closely with the product owner. Once I had an idea for the format, I created a list of questions we would want to ask, some of which were perhaps excessive, but again, I assumed more was better because in the moment I could trim the list by the way the customer answered questions, but I might fail to think of a question which could not be 'included' unless I spent time thinking about it. Ironically this felt a little heavy (for a group of judges which is composed of lots of pro CDT people), but again, these sorts of reports are heavy in my view.

At this point we had generically pre-processed the report and the interaction with the judges. We were told the day of that it was visual screen capture software so I did some research before the start of the contest, to compare feature sets. I captured some possible tests we might do and documented those.

While I have built my own screen recording tools, and even automation tools I have rarely gotten to the chance to test out such a massively well known support tool as SnagIt. We didn't know it was SnagIt until half an hour before the competition had actually started. With half an hour to go, we started to organize our thoughts and take a look around the tool. I create a few product oriented questions and found some obvious bugs. One the really concerned me was the order page link did not work. That turned out to be important later. I wrote notes down, but didn't file bugs as the contest had not started. I was just trying to understand the product.

We were told we could start asking questions, and I started peppering the judge with questions, but try to keep them slow enough to know if I need a follow up question or not. The judge's ability to respond to feedback was a little poor. For example I wrote a question regarding screen capture in Firefox which includes a plugin, something that seemed poorly documented. I missed saying it was for Mac as the plugin is only for Mac. They got confused by this and I updated my response but they were not looking at the chat log. I don't mind the video chat, but that became frustrating for me.

We as group decided two people would be in Mac land, one would be in windows VM and I would be between my personal win 8 box and my Mac.  We split up the functionality as well and started to test.  We did a load test using static fuzzing for a recording (E.G. something that wouldn't compress well) and my win 8 box has very little resources, so I did a load test where I recorded video while compiling.  I know lots of other testing was done, but those were some of the more interesting tests that I can recall (this was all written weeks after the contest was over).

As a tool smith who specialized in Windows, I appreciate how hard it is to build complex tools. However, when the product owner said that they used the tool to capture screenshots of bugs in the tool it seemed funny as we wrote up several bugs around that. We found that claim to be incorrect, but I wonder if the judges heard/recalled that claim since I didn't create strong paper-trail around it. Also because the judges in general don't have conversations with people afterwards it is hard to know what someone is thinking. I get that writing has to be clear, but how much context does one normally capture? Without knowing the culture, that might have been a throw away line from the product owner or it might have been serious. It makes me appreciate consultants in some ways, even if I do think they sometimes are more about conning and insulting than providing useful data.

It was genuinely neat to try to test something I would not have tested otherwise and doubt I will ever personally work on. It was fun. We found a lot of bugs. With one hour left I asked for permission to load test the download link, as that seemed part of the system and non-functional. We couldn't get a response. The judges were joking and talking about their experiences. Isaac tweeted two of them and they ignored them. The judges were filling the dead space that should have been silent with conversation, which I think was a detriment to the experience, but on the other hand, maybe that is part of the contest. Simulating that annoying coworker who keeps talking and talking in one long unbroken sentence moving from topic to topic... Yeah, Wayne! Quit that.

While finalizing our report, the judges sent a response to one of my early findings about the order link being broken saying that was expected.  I had put in a post 2+ hours earlier about that bug and the judges had not answered.  I didn't find out they had responded until after the contest was over.  We were given a chance to amend our report because of this, which I appreciate, but I also think it might have changed our testing strategy.  When you find one bug, you often look for clusters of bugs, which is what I did.  Obviously life is not fair and I don't expect the first year of a contest to run perfectly smoothly.

At this point we know our placing but not our scores. The contest was broken up by judge with points and scores attached. We got bonus points for helping one user, or at least that is what we were told. I wonder if we will get them or if their will be fear that giving said scores might cause disputes like you forgot to give points for X. It would suck if our placement went up or down at this point. But we shall see.  I will post a second blog entry based upon the score information we are given.

Friday, April 18, 2014

A Personal Failure Analysis: Why it appears CDT has failed

Talks & Failure Analysis

For the last few weeks I have been doing one presentation a week to various monthly groups and to a yearly event.  I have one more planned presentation in June, but this really isn't about my presentations, but rather the audience and fellow speakers.  

The first presentation was a meetup group Isaac and I helped form up in March.  Isaac and I gave the talk and it went really well.  Our audience was maybe 40% high-level QA with a few junior and the rest mid levels.  The talk was about testing an object in black box style.  We gave this talk 3 times total, and each time it seemed to go a little worse.  

When we did it with a group of developers, we moved to more white box/boundary testing style.  We had one person refuse to test out of principle.  In part I think us changing the style was at fault, but two sharper developers really got it, so I don't think it was totally our fault.  Two people really didn't get it until I changed boundary testing into a UI exercise.  This tells me even (junior) developers have a hard time visualizing what code does in their head.  

The final audience was a mixture of testers and developers.  The developers were often dis-interested in our talk and several of them left to go see a different track.  Either they felt applying black box techniques was not part of their domain or that they already knew this and wanted something new.  The irony of that is that 60% of the sessions were marked beginner, so learning something 'new' for a mid/senior developer is challenging at best.

I had a second talk which was a shorter version of my intended lecture I will be giving out of the country.  It was around reflections, something I think I have some level of mastery around, so I felt very prepared for my talk.  I will admit my cadence was a bit off, I probably rushed a little too much and perhaps didn't give enough audience breaks, but it was a code oriented talk.  Given the complexity, that might have been a small part of the problem but in reality, most of the audience were either junior developers or students who in their 2nd-4th year of college.  I had marked my talk as advanced but I couldn't control the people who attended.  This caused me to wonder something.  Why is it we are still talking about TDD, Intro to jquery, etc.?  I don't mean to say we shouldn't teach new people, but why are we mostly limited to beginners?

Why do we retrain people all the time?

In talking with 'experts' in the business, it becomes clear that my question of why do we repeat ourselves is because in some communities, such as testing, people don't stick with it in the long run.  Roughly 5 years appears to be the mean time between starting a career in testing and going to something else.  The second issue is that many people are not particularly interested in expanding their knowledge.  They just want to go to work, do their job and go home.  Don't get me wrong, I don't mean to say you should work 90 hour weeks, but many people have little passion for their work or improving their craft.  

Sometimes managers make these engineers look into 'new' development techniques, which create such a long tail into training that something like TDD continues to need talks 10 years later.  

With that analysis, I want to talk about it with an 'expert', and so I found the keynote speaker whom I spoke to at length.  I don't want to cite his name as I have not asked his permission nor do I have an easily accessible method to communicate with him (I don't have his email address and don't use twitter).  

He is a well-known developer who has a pod cast program with hundreds of different interviews with various people within the software industry.  In our discussion, I asked him if he knew Matt Heusser and he said he didn't.  I asked if he knew of James Bach, Cem Kaner, Rex Black...  He knew of no one I asked of, although having reviewed his interviews, the only tester I recognized that he interviewed was James Whittaker, which was many years ago.  I had not asked if he knew Whittaker.  I did ask if he knew any US experts/thinkers in test, but he declined to answer my question, saying he didn't think that way.  Perhaps my question was somehow confusing.

In speaking with this gentleman, a man who clearly had a wide variety of experience, and seemed to subscribe to the Analytical School of Testing, it became clear to me.  One of the reasons we end up not keeping testers around of long is because we as a group are not well known.  James Bach, in spite or perhaps because of his controversy, has one of the biggest names in testing.  Sadly, managers aren't demanding people start doing 'Bachian testing' (even though I suspect Bach himself would object to that name and concept).  We don't have a glamorous single one size fits all solution, like Scrum or TDD.  CDT is so vague no one but its practitioners knows exactly what it is.  It isn't easily measurable like Scrum or TDD which I can say if we are doing it or not (Are there sprints?  Are there unit tests?).  It makes me wonder if CDT has lost its war when managers can't easily measure if there CDT testers are even doing CDT testing, much less if CDT is working for them.  The metrics (be it sprints or number of unit tests) might ultimately not matter, but managers still at least one security blanket before they are likely to invest in structural change.

Solutions?

I'm not sure I have any useful solutions other than to say this appears to be a community issue.  I have tried doing mentor-ships with mild success.  It however doesn't scale beyond 1-2 people at a time.  I have tried doing lectures, but as soon as I stray past beginner style thinking or just talking about the concepts without talking about implementation, it seems to be clunky at best.  I have started a group with Isaac to see if a more personal monthly audience will help.  Maybe there we can make headway, but I don't know.  One nice thing with it being monthly and having multiple speakers, I might find out if the issue is in fact somehow with me and how I am presenting.  Even outside of presenting, maybe the way I communicate poorly and that just needs improvement?  For that matter, maybe I'm using the wrong language (English) to reach my audience of highly-motivate testers and software developers?  Perhaps there are more motivated testers outside of the US and my issues only revolve around where I live?

Perhaps I need to accept that I can only speak to a handful of experts, but how do I find and group them together?  How can the experts then package this stuff up together in such a way that they might learn from my knowledge?

Outside of myself, I know Kaner is moving to a more teaching oriented approach generating classes in which the students come to him.  Having taken a BBST course and talking to others who have, I know only about 1 in 3 (or less) of those students seem to both try and have the skills to succeed.  Some choose only to pass by the skin of their teeth and others simply don't get it.  Maybe some of it is because the course is in English, but it certainly isn't the only reason.  However, 1 in 3 might be as good as it ever gets, perhaps my expectations are just too high.  Kaner appears to have given up on a education only system by adding in certifications.  It appears to me the ultimate answer maybe that one has to market one's ideas more in order to convince people to change.  Perhaps the reason Kaner is talking about moving to such a path is because he failed to market his solution well enough.  Going back the keynote speaker I spoke to after my talk, he suggested that if he didn't know about the thinkers/experts of test or the ideas of CDT, they must not be that important.  Because if you don't hear about it, it can't matter.

Thursday, March 20, 2014

What is a comment? A miserable little pile of words!

A few weeks ago I tried an experiment.  I tried to comment on at least one test-related article or comment every day of the week.  I tried this for about a week and a half.  It was tiring, but interesting, with a result that made me sad.  I didn't do wimpy non-comments like "Great job".  I did in depth, well considered comments, that often took me an hour to develop.  And I got nothing back.  It gives me a new respect for people like James Bach who seem to be all over the Internet commenting all the time and doesn't seem to feel defeated.  As a consultant and trainer, that is a method of advertisement, but I do believe James Bach genuinely cares too.  However, my concerns with comments were developed much earlier...

I am a big fan of  Code Horror by Jeff Atwood who influenced my choice to do a blog in the first place.  While I don't always agree with him, I love his views on commenting.  In fact, he and James had a 'chat' about comments some years ago.  To summarize, James Bach had noticed a major flaw in Jeff Atwood's blog post and wrote his own post on it.  Jeff wanted to continue the conversation but couldn't because James didn't allow for comments.  Now that was years ago and once again, to Bach's credit, comments are enabled.  Victory.  We are all done.  Except...

Cranky Old Man


I still have complaints about how our community uses comments.  Certainly we have useless 'great job' comments without any critical thinking involved.  You can see them in even some of the bigger venues.  Even if you think the blog post is great, and you're their new biggest fan, that sort of comment helps no one.  I won't think about you, my biggest fan.  If you hate my work and yell at me without any particular logic, you're no better than the old man to the left.  However, as a blog writer, am I just a cranky old man yelling at human nature?

I don't think so.  Often Slashdot's comments are valuable and well thought out.  Bach's comments section is curated, often insightful and from some reports perhaps even censoring, but that is a different issue.  I think there are a few issues with our community.  The first is that we write too much.  You heard me.  An author says we write too much.  Ha!  I can read about 5x more than I can write.  When Markus Gartner on a nearly daily basis writes a 1-2 page blog, I think it is nearly impossible to have a conversation about any topic he brings up.  In two days Markus is off into the next subject.  Being 12 hours away doesn't help the conversation, but in the world of twitter, where detail doesn't matter but lip service is paid, you can't develop consensus or new ideas.

Why can't we slow down and as a community genuinely talk through issues?  I am amazed at those whom write daily, and I get that people like Matt Heusser who suggest writing at least 2000 words twice a week.  It isn't an invalid idea for improving writing, but that doesn't mean everything you write is worth posting.  I currently have some 2120 [this article is now published!] articles in draft format, including this one.  Not all of them make it out to you, the public.  I want to give you quality work, and hope to get quality comments back.  I don't write too much, in the hopes that I can learn from the community as I write articles.

But Wait...

What about all those spammers....?

Well to be honest, I rarely see them here.  I think Jeff Atwood might have a solution for bigger blogs, but I am reserving judgment for now.  Truthfully, that would be a good problem to have, as that would mean there are a lot more people caring about QA. 

What if I don't have anything to say?

Really?  You read a 1-2 page article, and found nothing of value to say? Nothing was added to your mental model?  You didn't see any mistakes, any logical leaps too far nor anything understated?  The blog post was so dull it would have been better off being ignored?  Which leads me to...

What if what I have to say is too mean?

Well, for this blogger, I give you permission to write mean things about my, JCD's blog posts.  In fact I WANT you to.  Please, say you hate it and then justify that opinion with facts.

...and if you do write a comment, with insight, which pushes me to write a response, I will try very hard to reply if not write an entire additional blog post on the subject.  So please do!

* Title with apologies to André Malraux.