17 January 2012

Challenge Analysis

I originally published this article on 1 November 2004.

Challenge Analysis


For the task of researching usability, the analyst has at their disposal an enormous range of tests.

From questionnaires and interviews, through protocol analysis, to performance data, OSM’s (and other forms of cognitive task analysis) one would expect all these things to be sufficient to measure everything that she wanted to know about a user.

But it still remains a difficult task to elucidate exactly what a user is thinking. Protocol analysis (Simon and Eriksson, 1982), or PA as I shall call it here, is also known as the “think aloud” analysis, for the user is sat in front of a system and asked to operate it while “thinking out aloud". Ideally, the user gets so used to doing this, that their cognition becomes easier to access. However, although it is a fantastic method, it does have drawbacks: in depends upon the user being a good verbaliser, and it can be slow to test a number of people. For these reasons, many tests only measure a few people. Other measures like questionnaires can be simple to issue and get completed, but the grain of analysis is often coarse.

However, one possibly new method occurred to me today, as an extension of protocol analysis, if you will. Maybe it’s time for analysts to get proactive!

Challenge Analysis.

The role of the analyst during PA is to remain quiet and let the user speak (with occasional prompts to remind them to verbalise their thoughts as much as they can). This is very much in the tradition of anthropological investigation, where the investigator aims to influence the subject of study as little as possible.

However, scientific experiment is all about intervention: without this, we would have no such thing as an experiment, rather we would only have observational studies. Clearly intervention is allowed, but it must be as tightly controlled as possible before it could be considered a worthwhile part of investigation.

My proposal is challenge analysis. This is very much like a PA, but it allows the analyst to intervene and ask questions of the user directly. While this is commonly done after PA using an interview technique, the user will often have forgotten the salient factors behind their decisions or behaviour. Interrupting them during the analysis and asking them to justify their actions could therefore be ruinous to a study if done incorrectly.

However, if done properly, a good challenge could extend the utility of PA somewhat.

Here’s an example. A user is writing a document using a text editor or a word processor. The user mentions that they need to move a block of text, and to do so they highlight it, copy it, move the cursor to the destination, paste it, move back to the original position, highlight the text again, and then delete it. In terms of actions, this takes longer than using the “cut” mechanism.

Observing this, I (the analyst) interrupt their work and ask them why they made this decision. If I am not satisfied, I can (diplomatically) illustrate any problems with their reasons and ask them again to justify what they are doing / what they were thinking, and what made them make that decision in the first place.

The disadvantages are many though.


  • It could create a confrontational atmosphere between the user and the analyst;
  • It might interrupt the user during the middle of a task causing them to alter their behaviour in a way they normally would not have;
  • It could provide the user with a new strategy which they will use in the future instead of their old strategy.
  • It stops the study being an observation study, and the analysts own personality may become too obvious;
  • The analyst would have to decide where and when to challenge (and when not to) - more experimenter bias;


However, a challenge analysis would also offer some advantages:


  • The user’s response to a challenge will more likely be accurate than that achieved from a post-PA interview;
  • Interesting or illustrative user behaviour that occurs during the study is less likely to be missed by the analyst;
  • Because the analyst may gain more insight into the users behaviour, the user may too - and come away from the study knowing more than when they arrived;

Clearly this is a very dangerous method to use, but with careful thought and good skills, both analyst and user may emerge from the experience richer in knowledge.

Outliers - a conundrum

I originally published this article on 28 October 2004 at Milui Articles

Outliers
a conundrum


Outliers are points of data that lie outside of what would be expected. They can be due to typographical errors (i.e., typing in an enormously incorrect number), participant error (i.e., forgetting to respond to a visual stimulus causing a very long reaction time), or due to an interesting effect that the researcher may not have considered before. However, calculating an outlier is may result in erroneous data being used for the proper analysis of research.

Consider these data: 1, 2, 3, 94. Of these four point, three of them are close together (the 1, 2 and 3), but the 94 value is way out of line with them. If a test showed these data together, the 94 could be considered an outlier because it is so different from the others.

In general terms, outliers are dealt with by either deleting them, transforming them, or investigating them further. What happens depends upon what type of outlier the data point is and what the researcher decides is the best way to deal with it.

Probably the most common process for determining outliers is to take the mean and a variance term of the data, and use these to examine which data points lie well outside the norm. A common method is to take the first quartile from the third (the interquartile range), and then calculating the boundaries by adding/subtracting them from the median. Anything lying between 1.5 - 3 times the median plus or minus this value may be considered a mild outlier, whereas anything more than the median plus or minus 3 times the interquartile range may be considered an extreme outlier.

A problem is however in working out what values should be used for these calculations. Using the first method described above (median +/- [1.5 * interquartile range]) is straightforward for univariate data, but what exactly goes into calculating the median and interquartile range? Should outliers be included into the calculation of the median and interquartile range which they are being compared to?

The rationale behind this idea is that an outlier (often) should not be there in the first place (definitely so with typographical errors), and is thus erroneous. Comparing data to a median and interquartile range that include erroneous values will produce an erroneous response, often in favour of the outlier (i.e., not identifying it). The GIGO (garbage in, garbage out) principle implies that any subsequent analyses are likely to be erroneous.

However, this is not true for all cases: if there are indeed no true outliers, then there is no problem as all the data are correct. However, if outliers are present, then we may have a problem. The best way to solve it is to exclude the outliers from the calculation of the median and interquartile range - that way, the erroneous data are being compared to correct values.

The conundrum in the title of this essay refers to this particular situation. If a researcher has a data set with outliers, they should do one thing, but if they don’t then they should do another. But the conundrum is that this decision cannot be made until the outliers have been identified.

In short, you need to know what the outliers are before you can discover that they are outliers!

How this could be resolved is discussed in a future article.

16 January 2012

Remote Testing - Or is it correct to test anonymous people over the Internet?


I originally published this article on Thursday 21 October 2004. This was before remote testing became commonplace so I like to indulge myself as being a bit ahead of the curve here.

Remote Testing

Or is it correct to test anonymous people over the Internet?

A tricky question indeed, and the cautious amongst us may say “no way!".

Why is that? Is there something about Internet users that automatically make them worthless as experimental participants? Not inherently, but the difficulty exists because it is impossible to verify that people are doing the test correctly.

If for example I wanted to test somebodies ability to navigate around a small website, how can I be sure that they haven’t done the test already and are redoing it (at a different computer and time) just to show that they can complete it to their satisfaction? The demand characteristics should be accounted for because they can confound the experiment.

From this there are three questions:


  1. Does this really happen?
  2. If it happens, is it of any statistical significance?
  3. If so, how can it be controlled?

The first question will depend largely upon the people who visit a website. A University with student-only parts may be able to ensure that they know exactly who is doing the test, but ‘out in the wild’, things are different. All it takes is a few jokers and the whole set of results would be worthless. Though I have no evidence for this, I expect my audience to be people who take this stuff more or less seriously: they are here because they are interested in HCI issues. If this is the case, then I think it is safe to assume that the people taking part in an online experiment can be trusted to be decent about it (but then I’m a hopeless optimist when it comes to human nature!).

The second question again depends upon the audience. As mentioned, if the first issue doesn’t arise, then this question (and the last one) are both moot which is good - go ahead and analyse the data. I would reckon that the best way forward would be to utilise some L-scores or something similar to test people. This though has the drawback that valid participants who fall outside the mean performance will be excluded (some might say that having an automatic exclusion for outliers would be a good thing, but I would rather examine the raw data first before making this decision).

The third question: how does one control for this? Again, this depends on the first two questions being issues. If not, then there is nothing to worry about, but if there is, the invalid participants have to be recognised and dealt with appropriately.

So what does this mean?

So now onto pragmatics.

Is there a way of testing whether a set of results are good or not? The basic idea within psychmetrics is the L score. This was designed to test responses by asking the same question from different points of view: the presence of inconsistencies indicate that the participant isn’t being entirely truthful. Two questions that could be used would be:


  • I am the life and soul of parties;
  • At busy social occasions, I prefer to stay quietly with the crowd.


Clearly, contradictory answers to these questions imply that the participant might (for example) be trying to answer positively to each question.

However, for a lot of cognitive psychology, these questions are hard to ask or incorporate into the design of an experiment: the designer must be cautious, or else the questions and their nature will stick out like a sore thumb, causing problems with the data. How would one ask the above two questions when one it trying to understand somebodies mental model of a web browser’s navigation system? Such questions must also not interfere with the testing itself: they must not provide cues to the answers of other questions. The above two examples, for instance, could easily apply to a measure of a persons extraversion. Possibly the best way is to ask questions central to the research aim but from different points of view. This may aid knowledge elicitation.

Statistical comparison may also be a viable method: outliers can be tested against the population average and disqualified if they lie outside the bounds of normality. This is often performed for many different analyses and is therefore valid, but the experimenter will need to be sure that the population average isn’t the basis of invalid responses (i.e., if you are testing 20 people and 4 of them are way out of line, can you be sure that the remaining 16 aren’t jokers?).

Following up a test with participants might be useful: contacting them at a later date (preferably in real time using chat or ICQ) to qualify their responses can often help the experimenter to estimate the participants’ likely intent (whether good or bad). In addition, this may also help the knowledge elicitation process as well.

A more rigorous solution would be to vet participants: test only those participants whose veracity can be ascertained beforehand. Of course, how this is done can be difficult, and it severely limits the utility of online testing by effectively reducing the test population significantly. However, this depends upon the design of the experiment and the experimenters wishes.

A final solution would be to alter the design of the experiment to increase its power: test more participants. The rationale behind this is that with more people tested, the invalid participants’ responses become less statistically significant during analysis. Indeed, in my own research , I have found that sometimes living with lots of variance can be possible within an empirical framework.

Summary.

In short, online testing can allow access to a larger population than could normally be tested. In addition, many people may be tested at once, reducing the workload of the experimenter significantly. However, in the way of many things, there are drawbacks in that the experimenter cannot know that the participant was actually tested properly. There are means of coping with this, but with online testing, the time and effort gains made will have to be set against the time and effort losses made in ensuring that the results are veracious.

Old articles from Milui

My first foray into UX freelance consultancy was through a company called Milui. The company is now defunct but I've managed to retrieve the articles and will be publishing them here. Article titles are:

Remote Testing
Outliers - a conundrum
Challenge Analysis
Practical Reliability of Experiments - a Practical Guide
Latent Semantic Analysis - an overview
So, you're short of cash and you need to test
Construct Validity
Give 'em what they want! Or should you?
Likert Scales and their Use
Why do people keep making "stupid" decisions?
Conversation Analysis and Collaborative Application Interfaces
The Learning Curve
Paper Review (Kane, 1994)
"The Principle of Genuine User Participation"
A New View of Research Validity Theory
Paper Review (Borsboom, Mellenbergh & van Heerden, 2004), The Concept of Validity
Paper Review: Downing (2003) Validity: on the Meaningful Interpretation of Assessment Data
How Many Items should go in a Menu?
Better Web Browser Usability?
Paper Review: Bowman & Hughes (2005)
So you have Writer's Block?
Problem Based Learning in Medical Education
Multiple Raters
Is "good enough" good enough?
The Rise and Rise of Search
Advanced Search
Designing for Amnesia
Web Usability - Non Relevant Links
Hate Dialogs, Love User Interaction?
50ms To Rate a Webpage!

I retrieved these articles from the WayBack machine and I'm excited to make them available again!

17 December 2011

Single-minded social networkers?

Some recent research of mine has been into social network and particularly how they encourage further participation and wider networks by suggesting people to 'follow', 'friend' or 'like'. Some of these are worrying me but not for the usual reasons of privacy invasion and so on...

My biggest concern is that the social graph is used to suggest others. This, in simplistic terms, tries to work out where a person (or 'node') lies in relation to every other person. Any other 'node' that lies near must be related, right? So if you follow someone, then that person's 'follows' must also be of interest to you.

But it has the implication that we are unidimensional. If one person who touches on a subject is of interest to us, then someone else who does the same must be? The assumption is that everyone is inter-connected and that a float number that indicates relatedness transfers to the real world. If I know person 'A' well and person 'A' knows 'person 'B' equally well, then I will probably know person 'B' well too. Except that it doesn't work like that. My colleagues might know my work self well but never have had the opportunity to meet my wife and vice versa. The 2 remain connected only through me; and then in different spheres of my life that may never coincide.

This idea is popular among engineers because it uses well-known algorithms to establish social proximity. It can be analysed, understood and relies on this assumption that is rarely questioned.

Admittedly, these are just suggestions and in no way are people forced into following total strangers; and these recommendations often do hit the target but they often fail too. And I believe that one of the causes is that recommendations based on the social graph are useful but only a part of the story.

Another part is the topic, the subjects about which a person writes. We write about what matters to us otherwise we wouldn't put the effort in. We write about things that are pertinent to our lives; irrelevant topics are not.

My argument is that relying on the social graph is good but to get closer to perfect recommendations, we need to use other ways to connect people, different types of information that can be used to find out how we relate to others, and therefore who is closely related to us as people.

One way we're doing this is at Roistr. Already, we're working on tying content together in a way that makes sense to people as a way to augment recommendations that can expand our social networks meaningfully.

11 December 2011

UX Research Principles I

I've spent a lot of time consulting for user research issues. My background and training have made research methods and statistics a real passion for me. I'm going to be writing a series of articles about how to undertake effective user research. This one is the first step: how to write an effective questions.
I like to think of myself as an experienced UX researcher. Being a psychologist, I've come across a lot of research methods that are not widely known but are very useful. This helps me to deal with problems better than if I didn't know them; and I've learned (the hard way!) how to use research methods to their best advantage.

But often in research, I've seen people diving straight in to decide upon the number of participants and research method. Being pro-active is commendable; but there is one important step that needs to be made before any of this other stuff can be done.

Set your questions.

It sounds obvious, so obvious that most people look puzzled when I say that we need to spend time nailing them down. But this comes from close on 15 years of research experience including a PhD and peer-reviewed publish articles detailing new research methods.

So why is it important to set the questions first? Well the questions determine everything else. They are the foundation of all research. With questions, you can then determine how they can be answered which implies what data you are going to collect and how to analyse it. When you have that, you can then decide upon a research method, then the target population, how to get your sample population, then your materials and then administer the lot. Finally, you can do the research.

What, step back a bit. All this comes from the questions?

Well yes. If you don't have the questions, how do you know what to measure? You cannot know what method you'll use or how many people you need. In that sequence, everything depends upon its preceding items. You cannot decide how many people to test unless you know what you're doing. A survey will need maybe 100 participants; user interviews maybe 6.

Some general pointers:

  • Take time over the questions and make sure they're good. If youve ever asked a statistician for advice, there's a good chance they'll ask you, "What's your research question?" or "What are you trying to find out about?"
  • Sit down with others and collaborate on formalising the research questions. Try to write them down like experimental hypotheses - something you can test.
  • Involving product managers, project managers, business owners, and other stakeholders can be a good for getting them to buy in to research. Just don't let it get bogged down with too much discussion; but ensure the questions are good.
  • Make a statement that summarises what the research is all about. This is good for explaining to people what is going on. The classic example mission statement is JFKs "by the end of the decade, we will put a man on the moon". A research statement might be something like, "What will make people use a social network site for cats?" If in doubt, always refer to this statement. If something isn't covered by it, then exclude it. This is why it's vital to work on making a good statement.
Once you have the questions established, you can work out how the answers to each question can be measured. More in the next article.

If you need any advice or help, leave a comment. If you have significant work on UX research that you think I can help with, you can contact me at alan@thoughtintodesign.com.

09 December 2011

Is a UX portfolio necessary?

Is a UX portfolio necessary? Should UX designers / researchers have a portfolio? IMHO, a UX portfolio is necessary but shouldn't be in an ideal world.



Some further reading links are below.

Currently, I've spent a lot of time putting together a web-based summary and a more complete print-based version to illustrate my process and the problems I've encountered. I've had to - recruiters have told me that it's expected and necessary for me to have a portfolio to be considered. The portfolio that got me interviews and jobs last year is now way too short.

But I also believe that a portfolio should not be necessary and there are X reasons why:

1) Documents cannot convey the true context of real-life work
2) A lot of UX thinking is hard to document thoroughly
3) People are after different levels of description
4) A lot of gatekeepers know little about UX
5) UX research & testing is hard to concisely communicate

Documents cannot convey the true context of real-life work

But from my 12 years experience in this field (and 16 years experience of web design), a portfolio only serves to gain immediate superficial attention. They rarely explain what the exact problems were solved and how they were challenged and met.

If a portfolio discussed a client and said, "This was the 5th time this project had been attempted: all past efforts had failed due to the clients' politics. I successfully pushed it through with 16 client stakeholders of a large, long-established multinational corporation most of whom cared most for building their empire rather than making the best quality product" it would sound bad particularly if the clients name was there despite being an excellent achievement in UX. Would I want to hire someone who bitches about their works' clients, even if they speak the truth? So this kind of critical information is hard to communicate.

A lot of UX thinking is hard to document thoroughly

I spend a lot of my work time thinking. I'll often go unapologetically to a coffee shop, sit down with a Moleskine and pencil and start jotting things down. Most of it's nonsense but I find even a change of situation to be conducive to creative thinking. I'd be short-changing the people who pay me money if I didn't do this.

What I'm trying to do is establish my own mental model (see Gentner & Stevens rather than Johnson-Laird and certainly not Young) about users' mental models - a supra-meta-mental model if you will! We all do it but to different degrees depending upon what information we have about users, what problems need to be solved and various other constraints.

There are artifacts from this thinking but they don't communicate well: Random sketches (with doodles inevitably thrown in) are the only really solid ones I have - they remain ideas until I can document them formally - and these documents don't show how I arrived at the solution, at least not without a long-winded (and frankly bizarre) explanation ("Well, I was reading a Mr Man book to my 3 year old daughter and one of the pictures of a house in a tree caught my attention. Somehow I put 2 and 2 together and realised that 'bit of information C' should go above 'bit of information A' and 'function B'." - this is really less than impressive even if it's honest).

Okay, we can write a corporate-friendly version ("I organised the UX team to ideate using concept-exploration and a semantic association task game") but this is nonsense and untrue.

People are after different levels of description

When reading about a problem, hirers are after different levels of explanation. It's a fact.

Some will read, "The problem was to design against a large set of business and legal constraints and I did this by paper wireframes, blah, blah, blah" and feel satisfied. After all, the candidate has shown that they have designed with business constraints in mind so that's good.

Personally (and there are others who feel the same), I feel this is superficial. What business and legal constraints? I want to know the detail before I can judge whether this work is good or not.

Example: When I worked for a bank, the business requirements were heavy: SOX compliance alone ensured that even short projects had 30 pages of constraints. The large ones - they were big documents. Designing against these was difficult because just taking the requirements in was an effort.

But I also cannot talk about the exact constraints because they are (or very well could be) corporate secrets. They're not for me to tell the outside world. This means that they have no place on a portfolio and hirers like myself would feel unsatisfied without this detail. This implies that portfolios are not the way to communicate UX ability.

To provide enough detail to make it worthwhile means including a lot of stuff that the hirer might not want to read. They're interest might be in navigating corporate politics; it might be in hammering through as good a design as possible against staid stakeholders; it might be whether they used ; or it might be 'are they easy to get on with?' Satisfying all information requirements will leave most overloaded. Not satisfying them results in meaningless communication.

Either the description is too long or there's just not enough detail.

A lot of gatekeepers know little about UX

Even a few people actually within UX don't know much about what UX really is.

But a lot of gatekeepers (HR, recruiters, creative background people) don't know much about what UX really is. It's not about graphic design; it's not about development or coding; it's not about business requirements (though all of these play a role). What differentiates UX from every other form of design is people:
knowing and finding out about how people think and behave and synthesising a solution that incorporates those user requirements.

It's why the word 'user' is in 'user experience'. If your work has no touch-points with users, then it's not UX.

Sadly, I'm encountering increasing numbers of people who just don't seem to realise this. Example: I was asked what my UX background was recently: "Are you from a design or development background?" The answer was neither: I'm a psychologist (though I can do limited graphic design and I can program to a surprising level - roistr.com is all my own work including the engine that drives it; and I wrote SalStat, a statistics program with a GUI back in the early 2000s including a lot of statistical tests - Mmmm, linear algebra!).

UX research & testing is hard to concisely communicate

I've done a lot of UX research and testing. My PhD is all research and even simple findings can be very hard to communicate to an expert audience never mind a naive one. Yeah it's our job but the research I do tends to be on a big scale. One I did recently featured something like 30-odd graphs.

However, a portfolio is a poor place to discuss this. This is because I cannot communicate the detail without releasing some very important corporate secrets. I can give you an anonymised version but it takes away all context and communicates nothing useful, at least nothing I'd be happy with.

Usability testing is similar. How can I communicate 10 hours of interviews and subsequent content analysis on the transcriptions? I'm good enough at my job to actually know what content analysis is and how to do it effectively (a rare skill in UX). But it's very challenging to communicate this: without context or detail, it's meaningless; with context and detail, it's breaking confidentiality. Besides, how does a hirer know that the work is of high quality? Very few in UX know anything about research anyway so there's little feedback coming through. I feel fairly confident in my work because I've been peer-reviewed and published and through a viva with a former Math Olympian (Alan Dix) and consultant to NASA (Andrew Howes). But few people from without a research background have had this training.

Conclusion

We need portfolios in order to get work. If we don't have them, jobs will go to those that do, simple as. It's not ideal.

But I would just be grateful if hirers could question the validity of basing decisions on portfolios; and also be aware of their own assumptions or even better, challenge them. Best would be to challenge themselves: being employed in UX doesn't qualify anyone as an expert.

I'm currently available for hire as a UX designer / researcher.

Now some links with discussions:

http://www.jeffgothelf.com/blog/you-dont-need-a-ux-portfolio/
http://www.quora.com/Is-a-portfolio-for-a-UX-Researcher-necessary
http://www.quora.com/Is-a-robust-portfolio-for-a-UX-Designer-necessary
http://www.uxmatters.com/mt/archives/2009/10/process-not-portfolio.php