17 January 2012

Multiple raters

This article was first published on 11 April 2005.

Multiple raters


So you have something that has been independently rated by several people. How can you tell if the ratings are the same? This article discusses how to approach the intraclass correlation coefficient in the area of inter-rater reliability.

Problem - are several ratings consistent?

Very often in usability research, an investigator will want to get an idea of how things are rated. Perhaps the investigator has been tasked with assessing the subjective readability of ten pages on a website and gets five people to offer ratings on a 5-point Likert scale. How then can he or she see whether the ratings are the same or not?

How then can he or she see whether the ratings are the same or not?

The best way to look at this is to use statistical analysis. Good analysis will allow the investigator to easily say whether there is a significant difference or not. However, I have commonly encountered people using correlations assuming that if all are significantly associated, then the ratings are the same. This is certainly a possible solution, but it’s tricky: using the above example, our investigator will have to perform ten different calculations: if any result is not statistically significant, then the ratings are not the same. In addition, by subjecting the same data to several, similar analyses, the investigator might be causing alpha inflation.. Because every analysis has a probability of one in twenty of happening due to chance, repeating analysis reduces this. With ten different tests, the probability of getting a significant result drops to one in two. The investigator would be making a serious error in doing this.

Intraclass correlations

But fear not! There are good tests that can be used to test the consistency of several raters in just one go. These are known as tests of inter-rater reliability.

Commonly, the statistic of interest is alpha, and the best test to use is the intraclass correlation. It’s available in newer versions of SPSS and was first discussed by Shrout & Fliess (1979) in the Psychological Bulletin. What this test does is very similar to performing several correlations all together, but in one test. This saves our investigator a lot of time (only one test to perform), the results are simpler, and there is no risk of alpha inflation.

However, not all is simple. There are three ways to perform an intraclass correlation and there are two statistics to use for each: six possibilities!

The choice of statistics depends upon whether the rating to be used will be assessed by one rater, or by more than one. In SPSS, these are referred to as single measures and average measures respectively.

The type of test to be used depends upon how the raters are selected. For the first type, each rater is selected at random from the population and rates only one case. For the second type, the raters judge every case, but the raters used are selected at random from the population. For the third type, raters judge every case, just the like second case. However, these are the only raters available.

For rules of thumb: if you have a set of raters and all of them are used, then use the third type of test (known as a fully-crossed 2-way anova mixed model design). Otherwise, you have the first or second type. The way to tell which one to use if only some of the possible raters are used is this: if the raters will judge only some cases, then you have a 1-way anova design and the first type should be used. The second type is to be used therefore if the raters judge every case (a fully-crossed 2-way anova design with random effects)

If you are using SPSS to perform an intraclass correlation, you will notice that 3 different statistics are given: the single measures, the group measures, and the alpha. Quite often, the alpha is the same as the group measures, but not always. The single measures is invariably lower than the average measures. Sometimes, results can be good (in terms of answering the research question) when the average measures statistic is used even if not appropriate whereas the single measures is too low. Don’t be tempted though to report the incorrect statistics: good peer review will ferret this out.

Other measures

Of course, there are other measures that can be used. The Kappa statistic can be calculated in many ways, but the intraclass correlation coefficient should be good enough to get you through most eventualities

References

Shrout, P. E., & Fleiss, J. L. (1979). Intraclass correlations: Uses in assessing rater reliability. Psychological Bulletin, 86, 420-428.

Problem based learning in medical education

This article was first published on 3 March 2005.

Problem based learning in medical education


Problem-based learning (PBL) is probably the largest area of educational research in medicine today. PBL is based firmly as a constructivist implementation whereby learning occurs through interaction with the environment and particularly other people. This essay considers the impact of PBL upon medical education with a short consideration of how PBL may be used in training novice users to handle computers.

In its broadest sense, PBL is implemented by presenting students with problems which they have to solve. It may be useful to draw a distinction between liberal implementations of PBL, and strict. The strict implementations require a firmly constructivist outlook: tutors should not “teach” per se, but rather be an interactive guide to the learning process. The whole idea of constructivism is that the learner is responsible for their own learning: the tutor simply ensures that an effective framework is set up.

Liberal implementations of PBL however simply rely upon the presentation of problems to encourage learning. These do not necessarily require an interaction between the learners and the instructors, and may be taught in an orthodox manner (using a problem as a worked example, for instance).

Most research considers the strict view to be the correct implementation, though there is no unanimity amongst researchers.

Norman and Schmidt (1992)

Norman and Schmidt (1992) reviewed the psychological basis of PBL, and said that Barrows (1986) held that PBL fosters clinical reasoning due to its closer approximation with clinical practice. The rationale behind PBL was that the acquisition of knowledge depends upon several mechanisms: the activation of prior knowledge, elaboration, and that matching content improves the recall of information. In a review of the literature, Norman and Schmidt found that PBL did not improve general problem solving ability. Gaps in knowledge were also more apparent, though it is not clear whether this was due to the way in which the learners were tested (wouldn’t testing learners’ knowledge with an orthodox method necessarily be higher if they were if taught using an orthodox method). However, the authors found that knowledge tended to be retained for longer, even though the amount of knowledge was less. PBL also enhanced interest in the topic matter as well as self-directed learning skills.

Albanese and Mitchell (1993)

Albanese and Mitchell (1993) performed a large scale meta-analysis of research papers on PBL from 1972 to 1992. Comparison between the papers was difficult due to the different methods involved, but some interesting findings were noted.

Broadly, the authors concluded that PBL was seen as more “nuturing” and “enjoyable” by the students when compared to traditional expository learning. In clinical exams and faculty interventions, graduates performed as well, if not better, than those taught with conventional methods.

However, they noted that the central difficulty with their study was defining exactly what PBL was: Barrows (1986) had defined a taxonomy of PBL types, saying PBL was characterised as the “use of patient problems as a context for students to acquire knowledge about the basic and clinical sciences.” The method involved encountering the problem, applying problem-solving with clinical reasoning skills, identifying learning needs in an interactive (often collaborative) process, self-study, applying newly gained knowledge to the problem, and then summarising it. In shorter terms, PBL may be defined as iteratively and interactively learning information needed to solve a problem. Within PBL, students are given a greater responsibility for the direction of their learning, though the exact amount differs across curricula. The main instructional method is through the use of small group tutorials, and independent study.

The theory behind PBL (from Schmidt, 1983) has three main points of action. Firstly prior knowledge should be activated. Though prior knowledge should not be at all sufficient to solve the problem, there should be enough to begin to make an approach. The instructional method must activate this knowledge.

Secondly, there should be encoding specificity. There should be a similar context between the instructional problem and commonly encountered cases.

Thirdly, there should be elaboration of knowledge through interactions such as peer-review, explanation, and interaction with others (learners and tutors).

Critics of conventional learning hold that the context of the learning differs from the context in which the learning is applied, so much so that a reduction in effectiveness is probable. Coles (1990) postulated that PBL equates more closely to contextual learning theory.

Albanese and Mitchell investigated five questions of interest to their study.

  1. The expense of PBL. Was any additional cost of PBL was worth changing curricula for?
  2. Effectiveness of self-directed learning. Do PBL learners learn as well as conventional students?
  3. Clinical skills may not depend upon the breadth of knowledge but rather on some other skill, so does PBL inhibit clinical skill?
  4. PBL may only prepare students for interactive small-group learning: real life situations may require more, so how do PBL learners adapt to other work environments?
  5. What time demands are placed upon staff?

The difficulties in measuring the outcomes of PBL were manifold. No “gold standard” for PBL assessment existed, so a range of other assessments had to be used to infer learning success or failure (such as standardised clinical achievement tests, clinical ratings of graduates in residency, self-ratings of background preparation, whether first choices of residency were achieved and so on).

An additional problem for investigating PBL was the selection for the groups: students who were placed into PBL groups were self-selected. When a conventional track was studied along with a PBL track, both tracks tended to blend, at least socially if not academically. Even without these issues, extraneous variables were not adequately controlled. However, the authors contend that the weight of consistent evidence should at the very least provide some support for their conclusions.

In terms of basic science knowledge, PBL students fared less well than conventional students. This appeared to be a general finding, though not wholly consistent as it depended upon the implementation of PBL. The more directive a PBL course was, the less likely was the difference.

For clinical examinations, there were few consistent differences. For one study, the PBL students appeared to be more homogenised, whereas the conventional students were dichotomously distributed. The general trend was in favour of PBL, though this was not conclusive.

Thought processes showed an advantage for PBL for atypical cases, whereas the reverse was true for classic cases. Errors were greater and decisiveness was lower for PBL. There was evidence that PBL may interfere with forward (expert) reasoning of clinical cases for PBL students tended to employ backward reasoning.

PBL students seemed to have more valued study behaviours: conventional students placed an emphasis upon using lecture notes and reproductive methods, whereas PBL students valued versatile methods. Use of libraries, textbooks, journals and article increased for PBL students.

Conventional students reported being more stressed than PBL students at the start of term. This difference lessened towards the end of the year, but was still significant. This was not a unanimous finding as one study reported greater stress for PBL students. However, this may be due to the specific implementation rather than the technique of PBL in general. PBL students’ least satisfactory aspects of their courses were the competition and essay examinations. Their most satisfactory aspects were problem solving, group discussions, applicability, and clinical relevance. Conventional students least liked the reliance on fact recall and multiple choice questionnaires. They preferred individual excellence and group competence. One study showed that PBL students viewed their preclinical years as engaging, difficult, and useful. Conventional students in this study viewed their preclinical years as irrelevant, passive, and boring.

For achieving the first choice of residency, PBL students compared well with conventional students.

Schmidt et al (1996)

Schmidt et al (1996) investigated medical courses held in three different institutions. One used conventional instruction, the second used problem-based learning, and the third had an integrative approach (whereby biomedical and clinical sciences were integrated around the major organ systems). The latter approach had structured elements (such as prescribed books), patient demonstrations and small group training sessions.

The theory behind PBL courses was that exposure to real life problems will enhance the problem-based craft because it fosters clinical reasoning and problem solving skills, though Norman and Schmidt (1992) did not find evidence for this.

This study could be criticised, for each instructional method took place at a different institution: effects of the quality of tutors and students could have confounded the research questions asked, and it does not appear as though these factors were adequately controlled. However, an interaction between instruction method and year of study was found, as well as main effects of both instruction method and year of study. The latter finding indicated some validity to the study: one would expect students to score more highly as they progressed through the course.

Further analysis of the main effect of instruction method showed that the PBL and integrative curricular did not differ between each other, whereas the conventional curricular achieved marks significantly lower than them both. The interaction effect showed that the integrative method achieved significantly higher marks than the PBL and conventional curricular for years 2 and 3. years 5 and 6 reverted to mirror the main effect by showing the PBL and conventional methods achieving higher marks than the conventional method, but not differing between each other.

The authors considered this to show that the benefits of PBL only become apparent during the clinical years of a students learning (and “incubation” period). During the clinical training however, the benefits become markedly apparent as performance rises to the level achieved by the integrative curricular. An alternative explanation is that the clerkship of the institution with the integrative method was better than the other two, but this was considered unlikely.

Human-computer interaction and user instruction

There is no reason why the benefits of PBL could not transfer to the area of user training: indeed, given that good interfaces are designed to allow exploration, it may be that a short series of problem-based exercises is the best way for users to be trained.

The one outstanding issue on this matter is whether experience using PBL as a user training technique will foster a more effective level of problem solving on computers.

One of the more intimidating aspects of seeking help for novice users is the ‘RTFM’ (”read the f***ing manual“!) and various other insults heaped upon the shoulders of those who dare to ask a question that the expert considers obvious. These experts justify their response by saying that the answers to simple questions may be found with a little searching on the part of the novice.

Though I have not looked at the literature for PBL within user training, I am not aware of any good studies that have been done. It may be that teaching users to search for solutions themselves would be a better solution than teaching them parrot-fashion where to click for a specific task. Indeed, given the amount of money spent on helpdesk operations, instituting a policy of PBL instead of conventional instruction may help to reduce the costs that organisations incur when dealing with emplyee’s IT problems.

Conclusion

Problem-based learning appears to offer very real benefits to learning. While all the above studies focused upon medical education, it is entirely likely that the benefits of PBL will transfer to other domains that have a practical aspect to them (such as architecture, psychology, and even human-computer interaction.

References

Albanese M, & Mitchell S (1993). Problem-based learning: A review of the literature on its outcomes and implementation issues. Academic Medicine. 68(1), 52-81.

Barrows H.S. & Tamblyn R.M. (1980) Problem-Based Learning: An Approach to Medical Education. New York: Springer Publishing Company, p.1.

Norman GR, Schmidt HG. The psychological basis of problem-based learning: A review of the evidence. Academic Medicine. 1992; 67: 557-565.

Schmidt, G., Machiels-Bongaerts, M., Hermans, H., Cate, Th.J. ten, Venekamp, R., & Boshuizen, H.P.A. (1996). The development of diagnostic competence: comparison of a problem-based, an integrated and a conventional medical curriculum. Academic Medicine. 71(6): 658-664.

Paper Review: Emotional responses of tutors and students in problem-based learning: lessons for staff development (Bowman & Hughes, 2005)


This article was first published on 31 January 2005.

Paper Review: Emotional responses of tutors and students in problem-based learning: lessons for staff development (Bowman & Hughes, 2005)

Bowman, D., and Hughes, P. (2005) Emotional responses of tutors and students in problem-based learning: lessons for staff development. Medical Education, 39, 145-153.

The authors detail a number of “emotional” risks of problem based learning and the possible group processes from which they may arise. Such processes can subvert the primary aim of PBL courses and lessen the effectiveness of how PBL might otherwise operate.

The group processes are explained with reference to psychotherapeutic groups: overrelianace on other group members, encouragement to the tutor to lower professional distance ("join the gang"), and secondary group concerns subverting the completion of the primary task. The authors conclude that attention should be paid to tutors’ emotional responses (anti-task behaviour).

Anti-task behaviour in tutors stems from five different areas:

When tutors act as therapists (tutors are not therapists);
When tutors want to “join in” with the students (they should have a professional distance from their students, i.e., be mentors and not friends);
When tutors try to keep control of the students (the risk of wanting to use their expert knowledge instead of letting the students guide themselves);
Subverting the primary task with a secondary one (often interpersonal conflicts);
Having a personal relationship with the students (again, professional distance is required);
The authors suggest solutions:

Agreement and clarity about tutors’ primary task (to prevent a secondary agenda);
Staff responsibilities have clear boundaries (prevent transgression by both parties);
Ongoing review and monitoring (reflection on group dynamics can help with current issues and bring awareness of potential problems);
Sociability that is not intimate (not distant, but professional);
Personally, I find the paper very conforming to PC standards. However, this is churlish for the aspect of maintaining a professional distance is laudable, as is insight into group processes and the possible risks that lie therein.

I would personally recommend that the findings of this paper are promulgated to tutors and faculty who use problem-based learning.

Better web browser usability?

This article was first published on 29 January 2005. It seems to have been implemented in IOS as 'Add to Reading List'.

Better web browser usability?


Scott Berkun has written a long article on how to build a better browser. There are some interesting links such as Abram’s empirical research into bookmarks [archive.org], and Tauscher’s work on history and navigation [also on archive.org]. I would also recommend Andy Cockburn’s work into people’s mental models of browser navigation mechanisms. Though this work is of a lower level than most browser designers need to know, it shows how even the programming operation of something should have HCI input to help prevent problems.

Scott discusses intelligent bookmarks, research and annotations, interaction with websites, and the like, and also discusses ideas that aren’t so good. His mention of intelligent bookmarks reminds me of something a colleague and friend of mine Hans Neth came up with - a temporary bookmark system that existed only for a session but was very malleable - he termed them anchors. However, the navigation of such a system was problematic: research has focused on ways of navigation using ancilliary controls (like maps), but users’ always seem to fall back on memory and the back button unless they are really lost.

Maybe that would be a good research question: how can you help provide a visual navigation system that doesn’t get in the way? Solutions with extra windows need not apply…!

How many items should go in a menu?

This article was first published on 29 January 2005.

How many items should go in a menu?


A lot of people think 7 ± 2 (i.e., between 5 and 9, with a preference for 7).

NO! It isn’t! And here I will explain why. But first I have to give a very short history lesson…

Miller (1956)

In 1956, Miller performed what has become a seminal experiment: he gauged the capacity of short term memory. What he found was that people could hold 7 ± 2 items in their memory (i.e., 5 to 9, tendency towards 7). This was astounding because nobody had before managed to define the limits of memory like this before.

Research expanded upon this somewhat (which is an understatement). For example, later research found evidence of “chunking” - people could remember more than 9 items if they managed to gather some items together into “chunks". Since then, an awful lot of work has been performed on it, and short term memory is still very much a hot topic in psychology.

But what has this to do with the number of items on a menu? Well, the thinking is that if a person can memorise 7 ± 2 items, then that is what should be in a menu. That’s the maximum amount that can be retained in short term memory without it being “processed” into long term memory.

Passing information into long term memory is a different process: consider short term memory as something like a cache: easily accessed, very fast, but just not there for long because there isn’t enough to go around. Long term memory by contrast is like saving something on a slow medium (like floppy disk, but many times larger). It’s slow, will probably take several tries, and it can be hard to get it back.

My brain hurts…

This sounds all well and good, but there are major problems. This is why:

If I perform a task using a computer, I have things in my short term memory. Maybe I want to copy a few words in a text editor and paste them somewhere else. I don’t want to commit this to long term memory: firstly it’s a waste (why would I want to remember this in several years time?), and secondly it’s slow. So I use my short term memory.

However, this is occupied with a few things already. Firstly what I am trying to do: what kind of document am I trying to create? Secondly, I want to remember what I want to move (and maybe I want to remember why as well). Thirdly, I want to remember where it is to be moved to (and maybe why as well). This list goes on.

You may have noticed that I need to keep these things in my memory. As I am sure you have experienced, you may forget why you are doing something in the middle of doing it. It’s awkward and slows you down, so it’s best to keep it “live” in the mind.

But with all these things floating around, surely my memory capacity is reduced? Well, put simply, yes it is. But if the theory behind 7 ± 2 items is correct, showing someone a menu with 9 items will cause them to entirely forget what they were doing, why they were doing it, and how. Is that really the case? It isn’t - I can look at a long menu and still remember exactly what it is that I was doing (in most cases!).

Recall versus recognition

Recall is when someone remembers something without any sensory prompt: recognition is when there is some prompt available.

For example, if I asked you in what year were you born, the chances are that you could recall it. However, if I asked you some obscure question about an API, you might not be able to tell me the answer ("it’s on the tip of my tongue!"), but if I gave you a hint (such as the name of a function) it would all come flooding back. That is recognition.

The 7 ± 2 items assumes that no long term memory is used when using a menu. However, menus rely on two different modes of cognition: recognition, and exploration.

As I explained, recognition relies on having a sensory prompt: the list of possible commands provides hints as to what can be done without needing to recall the exact command (which is why clearly defined menu actions are so important). With exploration, users can just guess at the best option and try it out, but that isn’t the issue here.

When people use menus, they rely on recognising the command: users who know the command but don’t know where it is can be seen searching for it, or examining the top level menu items for a hint. However, it is false to say that one “chunk” of memory is used for each menu item. I can look at a menu of 12 items and remember my original task.

I’ll look into the number of items in a future article.

References

Howes, A. (1994). A model of the acquisition of menu knowledge by exploration. In W. Kellog & T. Hewett (Eds.) Proceedings of the ACM Conference on Human Factors in Computing Systems CHI’94. Boston, MA. New York: ACM.

Miller, G. (1956). The magical number seven plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63, 81-97.

Review - Downing (2003) Validity: on the meaningful interpretation of assessment data

This article was first published on 27 January 2005.

Review - Downing (2003) Validity: on the meaningful interpretation of assessment data


Downing, S.M. (2003) Validity: on the meaningful interpretation of assessment data. Medical Education, 37, 830-837.

This paper concentrates largely upon the issue of validity and how it pertains to medical education. Downing says that all validity for medical education assessments is construct validity which can arise from five sources. Common misconceptions held about validity assessments are also discussed.

Validity should always be tested as a hypothesis, and is an essential component of research quality. Assessment tools are described as having degrees of validity rather than all or nothing properties. Curiously, this aspect of validity is somewhat at odds with the rest of the scientific method where a significant difference allows the researcher to say that “x affected y": never can a researcher say with confidence that their assessments are “valid” or not.

The description of validity is contained within the AERA Standards of Educational and Psychological Measurement, and now instead of there being different kinds of validity, all is now contained under a unitary concept, construct validity.

Construct validity arises from the postulation of theoretical constructs whose degree of relatedness to evidence is measured by empirical assessment.

Referring to the work of Kane, validity should be argued using an “evidentiary chain” with parallel means of investigation. This chain should related the empirical evidence with the “network of theory, hypotheses and logic". This process of analysis is an ongoing exercise.

Downing mentions that validity may be “typed” into five categories: content, response process, internal structure, relationship to other variables, and consequences. Within these sources are numbers of specific examples which Downing helpfully provides in a table.

In review, the concept of validity as discussed here is very much the modern interpretation of the traditional ideas of Cronbach, Messick and others. However, it seems a complex way to determine whether a “test measures what it is supposed to measure".

One problem I see with this concept (rather: collection of concepts) is that a lot of the assessment relies on correlational data: these cannot allow a researcher to make certain claims, for correlations only show associations. While this may not affect the results of an assessment in practical and immediate terms, one is still left with the ambiguity of the cause of observed effects. A later article by Borsboom et al explains all this in more detail, and also discusses the problem with the old diminished concept of criterion validity substituting itself for construct validity. The entire concept of validity as discussed here relies upon the construction of a workable nomological network of theory into which the proposed constructs can be placed. Much research into validity doesn’t bother with this, and instead reviews validity post hoc rather than assessing it within the bounds of the scientific method.

In summary, this paper appears to offer a concise view of validity that updates the traditional concepts that are commonly referred to in assessment work. However, I feel that serious problems remain with the entire concept of what validity is, how it can be assessed, and its importance to the process of empirical enquiry.

Review - Borsboom, Mellenbergh & van Heerden (2004) The concept of validity

This article was first published on 26 January 2005.

Review - Borsboom, Mellenbergh & van Heerden (2004) The concept of validity


Borsboom, D., Mellenbergh, G.J., & van Heerden, J. (2004) The concept of validity. Psychological Review, 111(4), 1061-1071.

This study begins by proposing a simplification in the concept of validity which is defined by two requirements: “a test is valid for measuring an attribute if (a) the attribute exists, and (b) variations in the attribute produce variation in the measurement outcomes” (in abstract). The authors make the claim that their theory is not only simpler but also theoretically superior. However, research purists may be concerned that the concept of validity decreases in scope and importance.

The concept of construct validity was expounded in a classic article by Cronbach & Meehl (1955) and expanded upon by Messick (1989), and the authors note that researchers’ conceptions of this work commonly differ. Very often, researchers will substitute the concepts of criterion validity for construct validity without being aware of the difference.

What seems to be the authors’ main criticism of construct validity is that research into the degree of a tests validity is based on epistemological characteristics, whereas effective measurement procedures in other fields use ontological claims. Ignoring the fact that ontology defines the effectiveness of a measurement scale means that the process of validation is difficult and tortuous (and it’s probably impossible to provide a clear answer).

Another objection to construct validity is its reliance on correlational measures and the way these are traditionally held to indicate the degree of validity of a test by having it compared to existing tests. The presence of a correlation (no matter how high) will not however infer any causative link. There are three grounds to this:

  1. criterion validity (which implied validity through the correlation of a test to a criterion) can imply that a test can measure many different things - if a series of universal characteristics affecting an experiment could be measured and compiled, there would be many correlations between variables and the test. The authors quote Guilforf (1946): “a test is valid for measuring many things". This is a weakness of criterion validity.
  2. the size of correlation equates with the degree of validity. The authors exemplify this by showing a perfect correlation between thunder and lightning - this does not though allow the measurement of lightning by measuring thunder. However, one response to this could be that both thunder and lightning arise from a common cause (an atmospheric event) which is something that the test does measure. But again, access to this event is not possible using associational measures.
  3. correlations are population-dependent statistics.

The authors propose their own definition which reduces the importance assigned to a tests validity. One interesting benefit is that validity becomes a dichotomous measure: either an attribute exists and the test measures it, or it does not.

The authors also feel that what was termed as experimental validity should be renamed, possibly as “overall quality". This would consist of validity, reliability (which is no longer a subset of validity), predictive accuracy, and absence of bias among others. One consequence of this is that validity would still be an essential component of research quality, but would be reduced in status.

Overall, the authors present what I feel is a compelling case. I not only feel that the current practices in validation are obscure and unworkable, but also overly complex. While some purists may not like the reduction of importance in validity, it does make sense.