Thursday, January 29, 2015

New algorithm can separate unstructured text into topics with high accuracy and reproducibility by Emily Ayshford

http://phys.org/news/2015-01-algorithm-unstructured-text-topics-high.html

Much of our reams of data sit in large databases of unstructured text. Finding insights among emails, text documents, and websites is extremely difficult unless we can search, characterize, and classify their text data in a meaningful way.

One of the leading big data algorithms for finding related topics within unstructured text (an area called topic modeling) is latent Dirichlet allocation (LDA). But when Northwestern University professor Luis Amaral set out to test LDA, he found that it was neither as accurate nor reproducible as a leading topic modeling algorithm should be.

Using his network analysis background, Amaral, professor of chemical and biological engineering in Northwestern's McCormick School of Engineering and Applied Science, developed a new topic modeling algorithm that has shown very high accuracy and reproducibility during tests. His results, published with co-author Konrad Kording, associate professor of physical medicine and rehabilitation, physiology, and applied mathematics at Northwestern, were published Jan. 29 in Physical Review X.

Topic modeling algorithms take unstructured text and find a set of topics that can be used to describe each document in the set. They are the workhorses of big data science, used as the foundation for recommendation systems, spam filtering, and digital image processing. The LDA topic modeling algorithm was developed in 2003 and has been widely used for academic research and for commercial applications, like search engines.

When Amaral explored how LDA worked, he found that the algorithm produced different results each time for the same set of data, and it often did so inaccurately. Amaral and his group tested LDA by running it on documents they created that were written in English, French, Spanish, and other languages. By doing this, they were able to prevent text overlap among documents.
"In this simple case, the algorithm should be able to perform at 100 percent accuracy and reproducibility," he said. But when LDA was used, it separated these documents into similar groups with only 90 percent accuracy and 80 percent reproducibility. "While these numbers may appear to be good, they are actually very poor, since they are for an exceedingly easy case," Amaral said.

To create a better algorithm, Amaral took a network approach. The result, called TopicMapping, begins by preprocessing data to replace words with their stem (so "star" and "stars" would be considered the same word). It then builds a network of connecting words and identifies a "community" of related words (just as one could look for communities of people in Facebook). The words within a given community define a topic.

The algorithm was able to perfectly separate the documents according to language and was able to reproduce its results. It also had high accuracy and reproducibility when separating 23,000 scientific papers and 1.2 million Wikipedia articles by topic.

These results show the need for more testing of big data algorithms and more research into making them more accurate and reproducible, Amaral said.

"Companies that make products must show that their products work," he said. "They must be certified. There is no such case for algorithms. We have a lot of uninformed consumers of big data algorithms that are using tools that haven't been tested for reproducibility and accuracy."

Friday, January 9, 2015

Top 10 Clever Google Search Tricks by Whitson Gordon

http://lifehacker.com/top-10-clever-google-search-tricks-1450186165


10. Use Google to Search Certain Sites

If you really like a web site but its search tool isn't very good, fret not—Google almost always does a better job, and you can use it to search that site with a simple operator. For example, if you want to find an old Lifehacker article, just type site:lifehacker.com before your search terms (e.g. site:lifehacker.com hackintosh). The same goes for your favorite forums, blogs, and even web services. In fact, it's actually really good for finding free audiobookssearching for free stuff without the spam, and more.

9. Find Product Names, Recipes, and More with Reverse Image Search

Google's reverse image search is great if you're looking for the source of a photo, wallpaper, or more images like that. However, reverse image search is also great for searching out information—like finding out who makes the chair in this picture, or how do I make the meal in this photo. Just punch in an image like you normally would, but look at Google's regular results instead of the image results—you'll probably find a lot.

8. Get "Wildcard" Suggestions Through Autocomplete

A lot of advanced search engines let you put a * in the middle of your terms to denote "anything." Google does too, but it doesn't always work the way you want. However, you can still get wildcard suggestions, of a sort, by typing in a full phrase in Google and then deleting the word you want to replace. For example, you can search for how to jailbreak an iphoneand remove one word to see all the suggestions for how to ____ an iphone.

7. Find Free Downloads of Any Type

Ever needed an old Android app but couldn't find the APK for what you were looking for? Or wanted an MP3 but couldn't find the right version? Google has a few search tools that, when used together, can unlock a plethora of downloads: inurlintitle, and filetype. For example, to find free Android APKs, you'd search for -inurl:htm -inurl:html intitle:"index of" apk to see site indexes of stored APK files. You can use this to find Android appsmusic filesfree ebookscomic books, and more. Check out the linked posts for more information.

6. Discover Alternatives to Popular Sites, Apps, and Products

You've probably searched for comparisons on Google before, like roku vs apple tv. But what if you don't know what you want to compare a product too, or you want to see what other competitors are out there? Just type in roku vs and see what Google's autocomplete adds. It'll most likely list the most popular competitors to the roku so you know what else to check out.You can also search for better than roku to see alternatives, too.

5. Access Google Cache Directly from the Search Bar

We all know Google Cache can be a great tool, but there's no need to search for the page and then hunt for that "Cached" link: just type cache: before that site's URL (e.g. cache:http://lifehacker.com). If Google has the site in its cache, it'll pull it right up for you. If you want to simplify the process even more, this bookmarklet is handy to have around. It's great for seeing an old version of a page, accessing a site when it's down, or getting past something like the SOPA blackout.

4. Bypass Paywalls, Blocked Sites, and More with a Google Proxy

You may already know that you can sometimes bypass paywalls, get around blocked sites, and download files by funneling a site through Google Translate or Google Mobilizer. That's a clever search trick in and of itself, but just like Google Cache, you can make the process a lot faster bykeeping a few URLs on hand. Just add the URL you want to visit to the end of the Google URL (e.g. http://translate.google.com/translate?sl=ja&tl=en&u=http://example.com/and you're good to go. Check out the full list of proxies, along with bookmarklets to make them even easier, here.

3. Search for People on Google Images

Some people's names are also real-world objects—like "Rose" or "Paris." If you're looking for a person and not a flower, just search for rose and add to &imgtype=facethe end of your search URL, as shown above. Google will redo the search but return results that it recognizes as faces!
Update: Reader unclghost kindly pointed out that we're working with outdated information here—this trick is now built into Google's UI! Just head to Search Tools > Type and you can choose from faces, photos, clip art, line drawings, and even animations. Thanks for the tip!

2. Get More Precise Time-Based Search Results

You've probably seen the option in Google that lets you filter results by time, such as the past hour, day, or week. But if you want something more specific—like in the past 10 minutes—you can do so with a URL hack. Just add &tbs=qdr: to the end of the URL, along with the time you want to search (which can include h5 for 5 hours, n5 for 5 minutes, or s5 for 5 seconds (substituting any number you want). So, to search within th past 10 minutes, you'd add&tbs=qdr:n10to your URL. It's handy for getting the most up-to-the-minute news.

1. Refine Your Search Terms with Advanced Operators

Okay, so this isn't so much a "clever use" than it is a tool everyone should have in their pocket. For everything Google can do, so few of us actually use the tools at our disposal. You probably already know you can search multiple terms with AND or OR, but have you ever used AROUND? AROUND is a halfway point between regular search terms (like white teeth) and using quotes (like "white teeth"). AROUND(2), for example, ensures that the two words are close to each other, but not necessarily in a specific order. You can tweak the range with a higher or lower number in the parentheses.
Similarly, if you want to exclude a word entirely, you can add a dash before it—like justin bieber -sucks if you want sites that only speak of Justin Bieber in a positive light. You can also use this to exclude other parameters—like excluding a site you don't like (troubleshooting mac -site:experts-exchange.com). Check out our guide to tweaking your Google searches for more of these tips, and you can also find a pretty solid list over at weblog Marc and Angel Hack Life. Search on!

Find In-Depth Articles on Google with a URL Trick by Whitson Gordon (works in America only)


If your Google search just isn't returning the quality content you want, this little URL trick might find more in-depth articles on the subject you're searching for.


Alex Chitu at Google Operating System recently discovered that Google has a section for "in-depth articles", from which it features longer posts from sites like the Wall Street Journal, New York Times, Wired, The Economist, and more. It only seems to work in the US, and it only pops up sometimes—but you can manually bring it up by adding this to the end of your search URL:
&tbs=ida:1&gl=us
It doesn't work all the time, and it's certainly a bit limiting, but it's worth a shot if Google just isn't giving you the kind of results you want. 

Wednesday, December 10, 2014

There is nothing new about the Knowledge Café or is there? by David Gurteen

https://www.linkedin.com/pulse/20141208122315-343667-there-is-nothing-new-about-the-knowledge-caf%C3%A9-or-is-there

There is nothing new about the Knowledge Café or is there? 

When people say that something is not new, they usually mean that they are familiar with the concept and its in common practice. 

To my mind, when this objection is levelled at the Knowledge Cafe - it means that they do not fully understand it. 

When I look at how organizations operate and the behaviours of people in organizations - it is quite apparent that people are either not aware of the fundamental principles and the power of good conversation or they understand them but do not to change their way of doing things either out of habit, laziness or choice. 

Why in meetings and presentations are we still so dependent on Powerpoint? Why is the dominant format of a talk, a long presentation with lots of Powerpoint slides and a very short time for Q&A? Why is no time included for reflection and no time for conversations amongst the participants in order for them to engage with the topic or issue? Why do we insist on talking at each other rather than with each other. 

Why is the dominant layout of our meeting rooms: either lecture style or large tables, when we know from experience and observation that these layouts are not conducive to good conversation? The research shows that good conversations take place in small groups of 3 or 4 people sitting around a small round table or even no table at all. 

Why in meetings, especially those where the people do not know each other well, do we not allow time for socialisation and relationship building before getting down to business when again the research shows that such socialisation improves people's cognitive skills. Why are circles rarely used in meeting's when the research and our own personal experience demonstrates their power? 

Why do managers and facilitators seek to control meetings so tightly and are afraid of negative talk or dissent. By surpressing people's fears, doubts and uncertainties - you do not eliminate them - you just drive them underground. Peter Block says "Yes" has no meaning if there is not the option to say "No". You need to bring people's doubts and fears out into the open and talk about them at length. 

And why when we know from research that group intelligence relates to how members of a team talk to each other. That it depends on the social sensitivity of the group members and on the readiness of the group to allow members to take equal turns in the conversation. And that groups where one person dominates are less collectively intelligent than in groups where the conversational turns are more evenly distributed, do we allow the same old people to dominate the conversations in our meetings and do nothing to encourage the quieter ones to engage and speak up. 

The Knowledge Cafe may not be totally new but it addresses all these issues and more but as a conversational method is still sadly very poorly adopted. 

In fact in many organizations conversation is seen as wasting time. But slowly this is changing. More and more people are starting to understand the power of conversation and take a conversational approach to the way that they connect, relate and work with each other. They see themselves as Conversational Leaders.

Wednesday, November 12, 2014

The Most Hilarious Proofreading Mistake in a Scientific Paper Ever by George Dvorsky

http://io9.com/the-most-hilarious-proofreading-mistake-in-a-scientific-1657839235


This is an actual quote from a scientific paper, published recently — and apparently without editing. Apparently the authors didn't think much of one of the papers they were citing. And their publisher didn't bother to edit out their pre-publication snark.
Ugh, this is not the kind of thing you want to see in a scientific journal. It makes us lose faith in peer review, and by consequence, the scientific method itself.
Four months after being published, someone finally noticed that a fish mating paper in the journal Ethology — "Variation in Melanism and Female Preference in Proximate but Ecologically Distinct Environments" — contained a rather embarrassing passage that both the authors and the peer reviewers failed to notice.
As Retraction Watch reported yesterday, the journal quickly removed the paper after the issue was brought to light.
Later, corresponding author Zach Culumber told Retraction Watch:
No, this was not intentional. It was added into the paper by a co-author during revision (after peer-review). It was unfortunately an oversight that became incorporated into the paper during the process of sending the manuscript back and forth between co-authors. The comment in question was not spotted during the proofing process with the journal. Neither myself nor any of the co-authors have any ill-will towards any other investigators, and I would never condone this sentiment towards another person or their work. We are working with the Journal now to correct the mistake. As the corresponding author, I apologize for the error.
Wiley says it's going to investigate the error and republish a corrected version as soon as possible, which now appears to have been done.

Sunday, November 9, 2014

Does Media Violence Predict Societal Violence? It Depends on What You Look at and When by Christopher J. Ferguson

http://onlinelibrary.wiley.com/doi/10.1111/jcom.12129/pdf

ABSTRACT:
This article presents 2 studies of the association of media violence rates with societal violence rates. In the first study, movie violence and homicide rates are examined across the 20th century and into the 21st (1920–2005). Throughout the mid-20th century small-to-moderate correlational relationships can be observed between movie violence and homicide rates in the United States. This trend reversed in the early and latter 20th century, with movie violence rates inversely related to homicide rates. In the second study, videogame violence consumption is examined against youth violence rates in the previous 2 decades. Videogame consumption is associated with a decline in youth violence rates. Results suggest that societal consumption of media violence is not predictive of increased societal violence rates.

Tuesday, October 21, 2014

Popular Mechanics: 6 Warning Signs That a Scientific Study is Bogus

http://www.popularmechanics.com/science/health/6-warning-signs-that-a-scientific-study-is-bogus-16674141

Was the Paper Published in a Peer-Reviewed Journal?


"If it wasn't, you have no reason to trust it," says Ivan Oransky, former executive editor at Reuters and cofounder of the blog Retraction Watch. "The peer-review system, as flawed as it is, stands between us and really poor science." Also, find out if the journal or its publisher is on Jeffrey Beall's list of questionable open-access journals, at scholarlyoa.com

What is the Journal's Impact Factor?


The impact factor is the average number of times a journal's papers are cited by other researchers. You can usually find this information on the journal's home page or by searching "impact factor" along with its name. Check out the impact factor of other journals in that field of research to see how they compare. 

Do the Researchers Cite Their Own Papers?


If so, this is a red flag that they are promoting views that fall outside the scientific consensus. Citations are listed at the end of a paper. 

How Many Test Subjects Were Used?


A large number of test subjects makes a study more robust and reduces the likelihood that the results are random. In general, the more questions a paper asks, the greater its sample size should be. Most reliable papers contain something called a p-value, which measures the probability (p) that a study's results occurred by random chance. In science a p-value of 0.05 suggests the study's conclusions may be meaningful. Smaller p-values are better. 

Does it Rely on Correlation?


Cigarette smoking has declined dramatically in the U.S. in the past few decades, and so has the national homicide rate. But just because two events occur at the same time doesn't mean that one caused the other. 

Have the Results Been Reproduced?


To find out, search the paper's name on Google Scholar and click on the Cited By link beneath the name. This will list other researchers who mention the paper in their own publications, and may also give you a clearer view of how other researchers critiqued the paper.