Now I'm going to discuss how we would look for a new law. In general, we look for a new law by the following process: first, we guess it, no, don’t laugh, that’s the truth. Then we compute the consequences of the guess, to see what, if this is right, if this law we guessed is right, to see what it would imply and then we compare the computation results to nature or we say compare to experiment or experience, compare it directly with observations to see if it works.
If it disagrees with experiment, it’s wrong! In that simple statement is the key to science. It doesn’t make any difference how beautiful your guess is, it doesn’t make a difference how smart you are, who made the guess, or what his name is… If it disagrees with experiment, it’s wrong. That’s all there is to it.
It is therefore not unscientific to make a guess, although many people who are not in science think it is. For instance, I had a conversation about flying saucers, some years ago, with a layman — because I am scientific I know all about flying saucers! I said “I don’t think there are flying saucers”. So the other...my antagonist said, “Is it impossible that there are flying saucers? Can you prove that it’s impossible?” “No, I can’t prove it’s impossible. It’s just very unlikely”. At that he said, “You are very unscientific. If you can’t prove it impossible then how can you say that it’s unlikely?” But that is the way that is scientific. It is scientific only to say what is more likely and what less likely, and not to be proving all the time the possible and impossible. To define what I mean, I might have said to him, "Listen, I mean that from my knowledge of the world that I see around me, I think that it is much more likely that the reports of flying saucers are the results of the known irrational characteristics of terrestrial intelligence than of the unknown rational efforts of extra-terrestrial intelligence." It is just more likely. That is all, and it is a very good guess. And we always try to guess the most likely explanation, keeping in the back of our minds the fact that if it does not work, then we must discuss the other possiblities.
There was, for instance, for a while, a phenomenon called super-conductivity, there still is the phenomenon, which is that metals conduct electricity without resistance at low temperatures and it was not at first obvious that this was a consequence of the known laws with these particles. Now that it has been thought through carefully enough, it is seen in fact to be fully explainable in terms of our present knowledge.
There are other phenomena, such as extra-sensory perception, which cannot be explained by our knowledge of physics here. However, that phenomenon has not been well established, and we cannot guarantee that it is there. If it could be demonstrated, of course, that would prove that physics is incomplete, and it is therefore extremely interesting to physicists whether it is right or wrong. Many, many experiments exist which show that it doesn't work. The same goes for astrological influences. If it were true that the stars could affect the day that it was good to go to the dentist - in America we have that kind of astrology - then the physics theory would be wrong, because there is no mechanism understandable in principle from these things that would make it go. That is the reason that there is some scepticism among scientists with regard to those ideas.
Now you see of course that with this method we can disprove any definite theory. We have a definite theory, a real guess, from which you can clearly compute consequences which could be compared to experiment and in principle we can get rid of any theory. You can always prove any definite theory wrong. Notice however that we never prove it right.
Suppose you invent a good guess, calculate the consequences, and discover every time that the consequences you have calculated agree with experiment. The theory is then right? No, it is simply not proved wrong.
Another thing I must point out is that you cannot prove a vague theory wrong. If the guess that you make is poorly expressed and rather vague, and the method that you use for figuring out the consequences is a little vague —you are not sure, and you say, “I think everything’s right because it’s all due to so and so, and such and such do this and that more or less, and I can sort of explain how this works...” then you see that this theory is good, because it cannot be proved wrong! Also if the process of computing the consequences is indefinite, then with a little skill any experimental results can be made to look like the expected consequences. You are probably familiar with that in other fields. ‘A’ hates his mother. The reason is, of course, because she did not caress him or love him enough when he was a child. But if you investigate you find out that as a matter of fact she did love him very much, and everything was all right. Well then, it was because she was overindulgent when he was a child! By having a vague theory it is possible to get either result. The cure for this one is the following: if it were possible to state exactly, ahead of time, how much love is not enough, and how much love is over-indulgent, then there would be a perfectly legitimate theory against which you could make tests. It is usually said when this is pointed out--when you are dealing with psychological matters things can’t be defined so precisely. Yes, but then you cannot claim to know anything about it.
Bernard Woolley: He's going to say something new and radical in the broadcast.
Sir Humphrey: What, that silly Grand Design? Bernard, that was precisely what you had to avoid! How did this come about, I shall need a very good explanation.
Bernard Woolley: Well, he's very keen on it.
Sir Humphrey: What's that got to do with it? Things don't happen just because Prime Ministers are very keen on them! Neville Chamberlain was very keen on peace.
Bernard Woolley: He thinks ... he thinks it’s a vote winner.
Sir Humphrey: Ah, that’s more serious. Sit down. What makes him think that?
Bernard Woolley: Well the party have had an opinion poll done and it seems all the voters are in favour of bringing back National Service.
Sir Humphrey: Well have another opinion poll done to show that they’re against bringing back National Service.
Bernard Woolley: They can’t be for and against
…
Sir Humphrey: Oh, of course they can Bernard! Have you ever been surveyed?
Bernard Woolley: Yes, well not me actually, my house … Oh I see what you mean
Sir Humphrey: You know what happens: nice young lady comes up to you. Obviously you want to create a good impression, you don’t want to look a fool, do you?
Bernard Woolley: No
Sir Humphrey: So she starts asking you some questions: Mr. Woolley, are you worried about the number of young people without jobs?
Bernard Woolley: Yes
Sir Humphrey Appleby: Are you worried about the rise in crime among teenagers?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Do you think there is lack of discipline in our Comprehensive Schools?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Do you think young people welcome some authority and leadership in their lives?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Do you think they respond to a challenge?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Would you be in favour of reintroducing National Service?
Bernard Woolley: Oh, well I suppose I might.
Sir Humphrey Appleby: Yes or no?
Bernard Woolley: Yes.
Sir Humphrey: Of course you
would, Bernard. After all you told you can’t say no to that. So they
don’t mention the first five questions and they publish the last one.
Bernard Woolley: Is that really what they do?
Sir Humphrey: Well, not the reputable ones, no, but there aren’t many of those. So alternatively the young lady can get the opposite result.
Bernard Woolley: How?
Sir Humphrey Appleby: Mr. Woolley, are you worried about the danger of war?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Are you worried about the growth of armaments?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Do you think there's a danger in giving young people guns and teaching them how to kill?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Do you think it's wrong to force people to take arms against their will?
Bernard Woolley: Yes.
Sir Humphrey Appleby: Would you oppose the reintroduction of National Service?
Bernard Woolley: Yes.
Sir Humphrey Appleby: There you are, you see, Bernard. The perfect balanced sample.
Differences and chance cause variation.
The real world varies unpredictably. Science is mostly about
discovering what causes the patterns we see. Why is it hotter this
decade than last? Why are there more birds in some areas than others?
There are many explanations for such trends, so the main challenge of
research is teasing apart the importance of the process of interest (for
example, the effect of climate change on bird populations) from the
innumerable other sources of variation (from widespread changes, such as
agricultural intensification and spread of invasive species, to
local-scale processes, such as the chance events that determine births
and deaths).
No measurement is exact.
Practically all measurements have some error. If the measurement
process were repeated, one might record a different result. In some
cases, the measurement error might be large compared with real
differences. Thus, if you are told that the economy grew by 0.13% last
month, there is a moderate chance that it may actually have shrunk.
Results should be presented with a precision that is appropriate for the
associated error, to avoid implying an unjustified degree of accuracy.
Bias is rife.
Experimental design or measuring devices may produce atypical results
in a given direction. For example, determining voting behaviour by
asking people on the street, at home or through the Internet will sample
different proportions of the population, and all may give different
results. Because studies that report 'statistically significant' results
are more likely to be written up and published, the scientific
literature tends to give an exaggerated picture of the magnitude of
problems or the effectiveness of solutions. An experiment might be
biased by expectations: participants provided with a treatment might
assume that they will experience a difference and so might behave
differently or report an effect. Researchers collecting the results can
be influenced by knowing who received treatment. The ideal experiment is
double-blind: neither the participants nor those collecting the data
know who received what. This might be straightforward in drug trials,
but it is impossible for many social studies. Confirmation bias arises
when scientists find evidence for a favoured theory and then become
insufficiently critical of their own results, or cease searching for
contrary evidence.
Bigger is usually better for sample size.
The average taken from a large number of observations will usually be
more informative than the average taken from a smaller number of
observations. That is, as we accumulate evidence, our knowledge
improves. This is especially important when studies are clouded by
substantial amounts of natural variation and measurement error. Thus,
the effectiveness of a drug treatment will vary naturally between
subjects. Its average efficacy can be more reliably and accurately
estimated from a trial with tens of thousands of participants than from
one with hundreds.
Correlation does not imply causation.
It is tempting to assume that one pattern causes another. However, the
correlation might be coincidental, or it might be a result of both
patterns being caused by a third factor — a 'confounding' or 'lurking'
variable. For example, ecologists at one time believed that poisonous
algae were killing fish in estuaries; it turned out that the algae grew
where fish died. The algae did not cause the deaths2.
Regression to the mean can mislead.
Extreme patterns in data are likely to be, at least in part, anomalies
attributable to chance or error. The next count is likely to be less
extreme. For example, if speed cameras are placed where there has been a
spate of accidents, any reduction in the accident rate cannot be
attributed to the camera; a reduction would probably have happened
anyway.
Extrapolating beyond the data is risky.
Patterns found within a given range do not necessarily apply outside
that range. Thus, it is very difficult to predict the response of
ecological systems to climate change, when the rate of change is faster
than has been experienced in the evolutionary history of existing
species, and when the weather extremes may be entirely new.
Beware the base-rate fallacy.
The ability of an imperfect test to identify a condition depends upon
the likelihood of that condition occurring (the base rate). For example,
a person might have a blood test that is '99% accurate' for a rare
disease and test positive, yet they might be unlikely to have the
disease. If 10,001 people have the test, of whom just one has the
disease, that person will almost certainly have a positive test, but so
too will a further 100 people (1%) even though they do not have the
disease. This type of calculation is valuable when considering any
screening procedure, say for terrorists at airports.
Controls are important. A control
group is dealt with in exactly the same way as the experimental group,
except that the treatment is not applied. Without a control, it is
difficult to determine whether a given treatment really had an effect.
The control helps researchers to be reasonably sure that there are no
confounding variables affecting the results. Sometimes people in trials
report positive outcomes because of the context or the person providing
the treatment, or even the colour of a tablet3.
This underlies the importance of comparing outcomes with a control,
such as a tablet without the active ingredient (a placebo).
Randomization avoids bias.
Experiments should, wherever possible, allocate individuals or groups
to interventions randomly. Comparing the educational achievement of
children whose parents adopt a health programme with that of children of
parents who do not is likely to suffer from bias (for example,
better-educated families might be more likely to join the programme). A
well-designed experiment would randomly select some parents to receive
the programme while others do not.
Seek replication, not pseudoreplication.
Results consistent across many studies, replicated on independent
populations, are more likely to be solid. The results of several such
experiments may be combined in a systematic review or a meta-analysis to
provide an overarching view of the topic with potentially much greater
statistical power than any of the individual studies. Applying an
intervention to several individuals in a group, say to a class of
children, might be misleading because the children will have many
features in common other than the intervention. The researchers might
make the mistake of 'pseudoreplication' if they generalize from these
children to a wider population that does not share the same
commonalities. Pseudoreplication leads to unwarranted faith in the
results. Pseudoreplication of studies on the abundance of cod in the
Grand Banks in Newfoundland, Canada, for example, contributed to the
collapse of what was once the largest cod fishery in the world4.
Scientists are human.
Scientists have a vested interest in promoting their work, often for
status and further research funding, although sometimes for direct
financial gain. This can lead to selective reporting of results and
occasionally, exaggeration. Peer review is not infallible: journal
editors might favour positive findings and newsworthiness. Multiple,
independent sources of evidence and replication are much more
convincing.
Significance is significant. Expressed as P, statistical significance is a measure of how likely a result is to occur by chance. Thus P
= 0.01 means there is a 1-in-100 probability that what looks like an
effect of the treatment could have occurred randomly, and in truth there
was no effect at all. Typically, scientists report results as
significant when the P-value of the test is less than 0.05 (1 in 20).
Separate no effect from non-significance. The lack of a statistically significant result (say a P-value
> 0.05) does not mean that there was no underlying effect: it means
that no effect was detected. A small study may not have the power to
detect a real difference. For example, tests of cotton and potato crops
that were genetically modified to produce a toxin to protect them from
damaging insects suggested that there were no adverse effects on
beneficial insects such as pollinators. Yet none of the experiments had
large enough sample sizes to detect impacts on beneficial species had
there been any5.
Effect size matters.
Small responses are less likely to be detected. A study with many
replicates might result in a statistically significant result but have a
small effect size (and so, perhaps, be unimportant). The importance of
an effect size is a biological, physical or social question, and not a
statistical one. In the 1990s, the editor of the US journal Epidemiology
asked authors to stop using statistical significance in submitted
manuscripts because authors were routinely misinterpreting the meaning
of significance tests, resulting in ineffective or misguided
recommendations for public-health policy6.
Study relevance limits generalizations.
The relevance of a study depends on how much the conditions under which
it is done resemble the conditions of the issue under consideration.
For example, there are limits to the generalizations that one can make
from animal or laboratory experiments to humans.
Feelings influence risk perception.
Broadly, risk can be thought of as the likelihood of an event occurring
in some time frame, multiplied by the consequences should the event
occur. People's risk perception is influenced disproportionately by many
things, including the rarity of the event, how much control they
believe they have, the adverseness of the outcomes, and whether the risk
is voluntarily or not. For example, people in the United States
underestimate the risks associated with having a handgun at home by
100-fold, and overestimate the risks of living close to a nuclear
reactor by 10-fold7.
Dependencies change the risks.
It is possible to calculate the consequences of individual events, such
as an extreme tide, heavy rainfall and key workers being absent.
However, if the events are interrelated, (for example a storm causes a
high tide, or heavy rain prevents workers from accessing the site) then
the probability of their co-occurrence is much higher than might be
expected8.
The assurance by credit-rating agencies that groups of subprime
mortgages had an exceedingly low risk of defaulting together was a major
element in the 2008 collapse of the credit markets.
Data can be dredged or cherry picked.
Evidence can be arranged to support one point of view. To interpret an
apparent association between consumption of yoghurt during pregnancy and
subsequent asthma in offspring9,
one would need to know whether the authors set out to test this sole
hypothesis, or happened across this finding in a huge data set. By
contrast, the evidence for the Higgs boson specifically accounted for
how hard researchers had to look for it — the 'look-elsewhere effect'.
The question to ask is: 'What am I not being told?'
Extreme measurements may mislead.
Any collation of measures (the effectiveness of a given school, say)
will show variability owing to differences in innate ability (teacher
competence), plus sampling (children might by chance be an atypical
sample with complications), plus bias (the school might be in an area
where people are unusually unhealthy), plus measurement error (outcomes
might be measured in different ways for different schools). However, the
resulting variation is typically interpreted only as differences in
innate ability, ignoring the other sources. This becomes problematic
with statements describing an extreme outcome ('the pass rate doubled')
or comparing the magnitude of the extreme with the mean ('the pass rate
in school x is three times the national average') or the range ('there is an x-fold
difference between the highest- and lowest-performing schools'). League
tables, in particular, are rarely reliable summaries of performance.
schweini writes "Inmates in an Oklahoma prison developed software that attempts to streamline the prison's food logistics.
A state representative found out, and he's trying to get every other
prison in Oklahoma to use it, too. According to the Washington Post,
'The program tracks inmates as they proceed through food lines, to make
sure they don’t go through the lines twice... It can help the prison
track how popular a particular meal is, so purchasers know how much food
to buy in the future. And it can track tools an inmate checks out to
perform their jobs.' The program also tracks supply shipments into the
system, and it showed that food supplier Sysco had been charging
different prices for the same food depending on which facility it was
going to. Another state representative was impressed, but realized the
need for oversight: 'If they build on what they’ve done here, they
actually have to script it out. If you have inmates writing code, there
has to be a continual auditing process. Food in prison is a commodity.
It’s currency.'"
There
are lots of ways to mess with the heads of undergraduate students.
Giving them a research assignment and failing to specify a minimum
number of references needed is just one example.
“Include as many sources as you need to make your point and illustrate your thesis.”
For students, finding one scholarly article on their topic often
seems to be enough. Researchers did an experiment, got some results,
and answered the research question the student started with. All done,
all set, time for dinner.
But science doesn’t work that way. One experiment may suggest
something interesting, but it doesn’t prove anything. In fact, it is
quite easy to point to many examples of intriguing scientific studies
that were either proved false or that couldn’t be reproduced later on.
Scientific ideas that are true should be reproducible: other researchers
should be able to repeat the experiments and get similar results or use
other methods to arrive at the same conclusions. You can’t say that
you discovered something new if someone else can’t reproduce your
result.
This fundamental scientific idea, reproducibility, may be in crisis.
A recent article by Vasilevsky et al. in the journal PeerJ suggested
that many scientific journal articles don’t provide the information that
other scientists would need in order to replicate their results. Key
information about chemicals, reactants or model organisms is often
missing, despite journal requirements to include such information
(Vasilevsky et al., 2013). And a recent item in The Economist
suggests that this might not matter that much. The emphasis placed on
new research (by funding agencies and tenure and promotion committees)
means that few scientists even attempt to replicate the work of others
(“Unreliable research: Trouble at the lab,” 2013).
All of this means trouble from the very beginning of a research
project, before an experiment is even designed, when scientists start to
do background research on their topics. In the same way that
experimental scientists can’t rely on the results of just one experiment
to prove something, relying on just one information source for
knowledge is a sure way to end up with unreliable information.
Journalists look for corroborating sources, wikipedia flags articles
that need a wider variety of citations, and scholars need to find
multiple scholarly articles to support their ideas.
Some innovative people, companies and publishers are trying to sort
this mess out. A collaboration between PLOS ONE, Mendeley, Figshare and
the Science Exchange will be attempting to replicate the results of
selected projects as a part of the Reproducibility Initiative. The Reproducibility Project is a crowdsourced effort to evaluate the reproducibility of experimental results in psychology. And the Reproducible Science
project aims to make the results of computational experiments
reproducible by ensuring the sharing of code and data and by making that
information available to reviewers who can test the results described
in a manuscript they are reviewing.
Unfortunately, these innovative programs are just a drop in the
bucket of modern science. Funding agencies, publishers and tenure and
promotion committees still value original work more highly than
verification work. Scientists who concentrated on replicating the work
of others would risk their careers.
As a result it is important for students and scholars to be aware of
the challenges facing the reproducibility of science. We teach students
in introductory science classes that reproducibility is one of the
hallmarks of science. As they learn more about their disciplines, they
need to be aware of the practical challenges involved in reproducing the
work of others, and the importance of finding multiple sources about a
topic needs to be emphasized.
As a librarian, part of my job is to help students find additional
sources related to their research topics, even if there isn’t a
published reproduction of an original source. This isn’t about which
database to use or whether to put quotes around a phrase. It is about
getting them to think critically about their topics. For example, while
there might not be a second study that repeated the experiment of the
first, students can look for:
Studies that examined the same topic in a different way
Studies that used the same methodology on a different species, geographic area, etc.
Background studies on individual aspects of their research question, including the statistical analyses used
Studies that cite the original study (even if no one has tried to
reproduce the results, other scholars might express doubts about their
conclusions when they cite the original).
The issues surrounding reproducibility in science won’t be solved
overnight, and it will take a concerted effort from scientists at all
levels of the modern scientific enterprise to steer this very big ship.
In the meantime, students and scholars can make special efforts to
ensure that they are using the highest quality information available as
the basis of their original studies.
Works Cited:
“Unreliable research: Trouble at the lab.” (2013, October 19). The Economist, 409(8858), 26-30.
Vasilevsky, N. a, Brush, M. H., Paddock, H., Ponting, L., Tripathy,
S. J., Larocca, G. M., & Haendel, M. A. (2013). On the
reproducibility of science: unique identification of research resources
in the biomedical literature. PeerJ, 1, e148. doi:10.7717/peerj.148.
One
of the trends we've seen is how, as the word of the NSA's spying has
spread, more and more ordinary people want to know how (or if) they can
defend themselves from surveillance online. But where to start?
The bad
news is: if you're being personally targeted by a powerful intelligence
agency like the NSA, it's very, very difficult to defend yourself. The
good news, if you can call it that, is that much of what the NSA is
doing is mass surveillance on everybody. With a few small steps, you can
make that kind of surveillance a lot more difficult and expensive, both
against you individually, and more generally against everyone.
Here's ten
steps you can take to make your own devices secure. This isn't a
complete list, and it won't make you completely safe from spying. But
every step you take will make you a little bit safer than average. And
it will make your attackers, whether they're the NSA or a local
criminal, have to work that much harder.
Use end-to-end encryption. We know the NSA has been working to undermineencryption, but experts like Bruce Schneier who have seen the NSA documents feel thatencryption is still "your friend".
And your best friends remain open source systems that don't share your
secret key with others, are open to examination by security experts, and
encrypt data all the way from one end of a conversation to the other:
from your device to the person you're chatting with. The easiest tool
that achieves this end-to-end encryption is off-the-record (OTR) messaging,
which gives instant messaging clients end-to-end encryption
capabilities (and you can use it over existing services, such as Google
Hangout and Facebook chat). Install it
on your own computers, and get your friends to install it too. When
you've done that, look into PGP–it's tricky to use, but used well it'll
stop your email from being an open book to snoopers.
Encrypt as much communications as you can. Even if you can't do end-to-end, you can still encrypt a lot of your Internet traffic. If you use EFF's HTTPS Everywhere
browser addon for Chrome or Firefox, you can maximise the amount of web
data you protect by forcing websites to encrypt webpages whenever
possible. Use a virtual private network (VPN) when you're on a network you don't trust, like a cybercafe.
Encrypt your hard drive. The latest version of Windows, Macs, iOS
and Android all have ways to encrypt your local storage. Turn it on.
Without it, anyone with a few minutes physical access to your computer,
tablet or smartphone can copy its contents, even if they don't have your
password.
Strong passwords, kept safe.
Passwords these days have to be ridiculously long to be safe against
crackers. That includes the password to email accounts, and passwords to
unlock devices, and passwords to web services. If it's bad to re-use
passwords, and bad to use short passwords, how can you remember them
all? Use a password manager.
Even write down your passwords and keeping them in your wallet is safer
than re-using the same short memorable password — at least you'll know
when your wallet is stolen. You can create a memorable strong master
password using a random word system like that described at diceware.com.
Use Tor."Tor Stinks",
this slide leaked from GCHQ says. That shows much the intelligence
services are worried about it. Tor is an the open source program that
protects your anonymity online by shuffling your data through a global
network of volunteer servers. If you install and use Tor,
you can hide your origins from corporate and mass surveillance. You'll
also be showing that Tor is used by everyone, not just the "terrorists"
that GCHQ claims.
Turn on two-factor (or two-step) authentication.Google and Gmail has it; Twitter has it; Dropbox
has it. Two factor authentication, where you type a password and a
regularly changed confirmation number, helps protect you from attacks on
web and cloud services. When available, turn it on for the services you
use. If it's not available, tell the company you want it.
Don't click on attachments.
The easiest ways to get intrusive malware onto your computer is through
your email, or through compromised websites. Browsers are getting
better at protecting you from the worst of the web, but files sent by
email or downloaded from the Net can still take complete control of your
computer. Get your friends to send you information in text; when they
send you a file, double-check it's really from them.
Keep software updated, and use anti-virus software.
The NSA may be attempting to compromise Internet companies (and we're
still waiting to see whether anti-virus companies deliberately ignore
government malware), but on the balance, it's still better to have the
companies trying to fix your software than have attackers be able to
exploit old bugs.
Keep extra secret information extra secure. Think about the data you have, and take extra steps to encrypt and conceal your most private data. You can use TrueCrypt
to separately encrypt a USB flash drive. You might even want to keep
your most private data on a cheap netbook, kept offline and only used
for the purposes of reading or editing documents.
Be an ally.
If you understand and care enough to have read this far, we need your
help. To really challenge the surveillance state, you need to teach
others what you've learned, and explain to them why it's important.
Install OTR, Tor and other software for worried colleagues, and teach
your friends how to use them. Explain to them the impact of the NSA
revelations. Ask them to sign up to Stop Watching Us and other campaigns against bulk spying. Run a Tor node, or hold a cryptoparty. They need to stop watching us; and we need to start making it much harder for them to get away with it.