Logo and Side Nav

News

Velit dreamcatcher cardigan anim, kitsch Godard occupy art party PBR. Ex cornhole mustache cliche. Anim proident accusamus tofu. Helvetica cillum labore quis magna, try-hard chia literally street art kale chips aliquip American Apparel.

Search

Browse News Archive

trends:

Showing posts with label trends. Show all posts
Showing posts with label trends. Show all posts

Monday, April 22, 2013

IEEE BIGDATA 2013: WORKSHOP ON BIG HUMANITIES


HIPerSpace_video_2


The Workshop on Big Humanities will be held in conjunction with the 2013 IEEE International Conference on Big Data (IEEE BigData 2013, 6-9 October 2013, San Francisco, California). The conference provides a leading international forum for disseminating the latest research in the growing field of “big data”.

The workshop will address applications of “big data” in the humanities, arts and culture, and the challenges and possibilities that such increased scale brings for scholarship in these areas.


Topics covered by the workshop include, but are not restricted to, the following:

Text- and data-mining of historical and archival material.
Social media analysis, including sentiment analysis
Cultural analytics
Crowd-sourcing and big data
Cyber-infrastructures for the humanities
Relationship between ‘small data’ and big data
NoSQL databases and their application, e.g. document and graph databases
Big data and the construction of memory and identity
Big data and archival practice
Construction of big data
Big data in Heritage



July 30, 2013: submission of full workshop papers

Other dates and submission details here.

Saturday, April 6, 2013

new PhD program in Digital Design at European Graduate School


Digital Design PhD Program at European Graduate School


Under its Director, Wolfgang Schirmacher, the European Graduate School [EGS] is currently inviting expressions of interest in a new 'Digital Design' stream within the EGS Media and Communication Division's Postgraduate Program leading to a possible PhD.

The EGS already has a very strong reputation in the area of Critical Theory, Philosophy and Media/Film Theory, counting some of the most illustrious thinkers in the world among its faculty, including Slavoj Zizek, Alain Badiou, Giorgio Agamben, Judith Butler, Helene Cixous, Jean Luc Nancy, Manuel De Landa, Lev Manovich, and Paul Virilio. It also has a strong reputation in the area of Media and Communications, with some of the leading figures in Film and Art. The EGS postgraduate program is designed for working professionals, and residential requirements are limited to attendance of a series of brief, intense workshops in the summer months at its base in Saas-Fee, Switzerland. For further details please see the EGS website: http://www.egs.edu/

The EGS is currently exploring the possibility of a new 'digital design' stream within the Media and Communication Division. If there prove to be sufficient initial expressions of interest from prospective students, the EGS will evaluate the establishment of this new program. The program will then begin in September 2013, with the first residential workshop taking place in Saas-Fee in June 2014, provided that it meets the basic minimum requirement of 20+ confirmed applicants.

The urge to establish a new Digital Design PhD stream within the Media and Communication Division at EGS stems in part from the limited number of PhD programs in this field, compared to the relative proliferation of masters programs around the world. It also stems from the growing expectation for academics to hold PhD degrees. But, above all, it stems from the demand for a relatively low cost PhD program that is flexible enough to allow those working in architectural offices or teaching in academic institutions to continue their employment while undertaking their PhD research. Current overall fees for a 4 year PhD with the EGS amount to $24, 450, making it an attractive financial option, compared to most other programs.

A number of pre-eminent figures from the field of Digital Design have already offered to teach on the new program:

John Frazer is regarded by many as the godfather of architectural computation, and acknowledged as a leader in the field of evolutionary digital design, and originator of the Evolutionary Digital Design Process. His seminal book, An Evolutionary Architecture, is regarded as a classic within the field.
www.johnfrazer.com/

Mark Burry is Professor of Innovation at RMIT, and Director of the Spatial Information Architecture Laboratory and founding Director of the Design Research Institute at RMIT. He is also executive architect and researcher for the Sagrada Familia Church in Barcelona, Spain.
http://www.sial.rmit.edu.au/People/mburry+Biography.php

Achim Menges is a Professor at the University of Stuttgart and a founding member of the Institute for Computational Design. His work focuses on the development of integral design processes at the intersection of morphogenetic computation, biomimetic engineering and computer aided manufacturing.
http://icd.uni-stuttgart.de/?p=897

Manuel De Landa is the Gilles Deleuze Professor of Contemporary Philosophy and Science at the European Graduate School. He also teaches at Pratt Institute, USC, and UPenn. He is the author of a series of highly influential books, including A Thousand Years of Non-Linear History and Intensive Science and Virtual Philosophy.
http://www.egs.edu/faculty/manuel-de-landa/biography/

Patrik Schumacher is a partner in Zaha Hadid Architects, and a founding director of the AA DRL. He studied both architecture and philosophy in London, Bonn and Stuttgart, and holds a PhD from Klagenfurt University. He has taught at many schools of architecture, and is the author of a two volume edition, The Autopoiesis of Architecture.
http://www.patrikschumacher.com/

Alisa Andrasek is an experimental practitioner and research based educator of architecture and computational processes in design. She now teaches design and theory seminars at the Bartlett School of Architecture, having previously taught at Columbia GSAP and the Architectural Association. She is a Director of Biothing/CONTINUUM.
http://www.biothing.org/?page_id=2

Mette Thomsen is Professor at the Royal Academy of Fine Arts in Copenhagen, where she directs the Centre for Information Technology and Architecture. Her work focuses on Digital Crafting as a way of thinking material practice, computation and fabrication as part of architectural culture.
http://cita.karch.dk/

Neil Leach is a designer and architectural theorist, who has published 23 books on architectural theory and digital design. He has taught at several leading schools of architecture, including SCI-Arc, Columbia GSAP, AA, DIA, IaaC and USC, and is a NASA Innovative Advanced Concepts Fellow developing robotic fabrication technologies for printing structures on the Moon and Mars.
http://arch.usc.edu/faculty/leach

In the first year instructors will each offer 3 days of intensive lectures over the residential summer workshop in Saas-Fee, Switzerland. These lectures are designed to fulfill the requirements of the European Credit Transfer System (ECTS) regarding contact hours and student workload.

During their second year students get the chance to study under the other illustrious EGS professors. Aside from Slavoj Zizek, and the other figures mentioned above, these also include others working in architecture and digital theory, such as Geert Lovink, Lev Manovich, Mitchell Joachim and Hendrik Speck. As they enter their third year, upon completion of all course work and examinations, students will be invited to establish supervisory arrangements with the individual professors. Research is expected to be self-directed according to the European model.

Students interested in the program are invited to send an initial expression of interest, together with a one page CV by 15 April 2013 to:


Neil Leach
leachneil@hotmail.com


Sunday, March 17, 2013

Call for Proposals: Graduate Center CUNY Provost’s Digital Innovation Grants



Visualization by Micki Kaufman, The Graduate Center CUNY
Visualization from the project by Micki Kaufman (PhD student in History, The Graduate Center CUNY): “Data Mining Diplomacy”: A Computational Analysis of the State Department’s Foreign Policy Files


Call for Proposals: Graduate Center CUNY Provost’s Digital Innovation Grants


Deadline: April 8, 2013


To submit: Send a single PDF file containing all parts of the application to gcdi@gc.cuny.edu with “Provost’s Digital Innovation Grant Proposal” in the subject line.

The Graduate Center CUNY Digital Initiatives project of the Provost’s Office is delighted to announce a call for proposals in support of innovative digital projects designed, created, programmed, or administered by matriculated doctoral students in good academic standing at the CUNY Graduate Center.

Proposal application details: http://gcdi.commons.gc.cuny.edu/2013/03/15/call-for-proposals-provosts-digital-innovation-grants/

Winning proposals from the 2012-2013 competition: http://gcdi.commons.gc.cuny.edu/category/provosts-digital-innovation-grants/

Award range: $500 to $3000


Friday, March 8, 2013

The Programmable City openings: 4 postdocs (5 years) and 4 PhD students (4 years)



code_space


The Programmable City

5 year research project directed by Rob Kitchin


Avaiable positions:

4 postdocs (5 years) and 4 PhD students (4 years)


The project is an empirical extension of the Code/Space book (Rob Kitchin and Martin Dodge) published in Software Studies series by The MIT Press (2011). It focuses on the intersection of smart urbanism, ubiquitous computing and big data from a software studies/critical geography perspective, comparing Dublin and Boston and other locales.

The positions are not restricted to any discipline.

Postdoctoral Researchers X 2 Posts (the othee 2 posts will be advertized later this year)
Closing date for applications 22nd March 2013
Further details

Funded PhDs X 4 Posts
Closing date for applications 12th April 2013
Further details



Prof. Rob Kitchin is Director of NIRSA, and Chairperson of the Irish Social Sciences Platform. He has published widely across the social sciences, including 20 books and over 100 articles and book chapters. He is editor of the international journals, Progress in Human Geography (ISI rank 2/61) and Dialogues in Human Geography, and for eleven years was the editor of Social and Cultural Geography. His book 'Code/Space' (with Martin Dodge) won the Association of American Geographers 'Meridian Book Award' for the outstanding book in the discipline in 2011 and a 'CHOICE Outstanding Academic Title 2011' award from the American Library Association.


Tuesday, November 27, 2012

the meaning of statistics and digital humanities


manga guide to statistics

"I still hear this persistent fear of people using computational analysis in the humanities bringing about scientism, or positivism. The specter of Cliometrics haunts us. This is completely backwards."
Trevor Owens, Discovery and Justification are Different: Notes on Science-ing the Humanities, 11/19/2012.


As the number of people using quantitative methods to study "cultural data" is gradually increasing (right now these people work in a few areas which do not interact: digital humanities, empirical film studies, computers and art history, computational social science), it is important to ask: what is statistics, and what does it mean to use statistical methods to study culture? Does using statistics immediately make you a positivist?

Here is one definition of statistics:

"A branch of mathematics dealing with the collection, analysis, interpretation, and presentation of masses of numerical data" (http://www.merriam-webster.com/dictionary/statistics)

Wikipedia article drops the reference to "mathematics" and repeats the rest:

"Statistics pertains to the collection, analysis, interpretation, and presentation of data." (http://en.wikipedia.org/wiki/Outline_of_statistics).

Without the reference to mathematics, this description looks very friendly - there is nothing here which directly calls for positivism, or scientific method.


But of course this is not enough to argue that statistics and humanities are compatible projects. So let's continue. It is standard to divide statistics into two approaches: descriptive and inferential.

"Descriptive statistics are used to describe the basic features of the data in a study. They provide simple summaries about the sample and the measures. Together with simple graphics analysis, they form the basis of virtually every quantitative analysis of data." (http://www.socialresearchmethods.net/kb/statdesc.php).

Examples of descriptive statistics are mean (the measure of central tendency) and standard deviation (the measure of dispersion).


Range_of_height_measurements_of_union_soldiers_1864
Adolphe Quetelet. A graph showing the distribution of height measurements of soldiers. This graph was a paet of many early statistical studies of Quetelet which led him to formulate a theory of "average man" (1835) which states that many measurements of human traits follow a normal curve. Source: E. B. Taylor, Quentelet on the Science of Man, Popular Science, Volume 1 (May 1872).


Application of descriptive statistics does not have to be followed by inferential statistics. The two serve different purposes. Descriptive statistics is only concerned with the data you have – it is a set of diverse techniques for summarizing the properties of this data in a compact form. In contrast, in inferential statistics, the collected data is only a tool for making statements about what is outside this sample (e.g, a population):

"With descriptive statistics you are simply describing what is or what the data shows. With inferential statistics, you are trying to reach conclusions that extend beyond the immediate data alone." (http://www.socialresearchmethods.net/kb/statdesc.php).


Traditionally, descriptive statistics usually summarized the data with numbers. Since the 1970, computers gradually made the use of graphs for studying the data (as opposed to only illustrating the findings) equally important. This was pioneered by John Tukey who came up with the term exploratory data analysis. "Exploratory data analysis (EDA) is an approach to analyzing data sets to summarize their main characteristics in easy-to-understand form, often with visual graphs, without using a statistical model or having formulated a hypothesis." (http://en.wikipedia.org/wiki/Exploratory_data_analysis). Tukey's work lead to the development of statistical and graphing software S, which in its turn lead to R, which today is the most widelly used computing platform for data exploration and analysis.

Accordingly, the current explanation of descriptive statistics on Wikipedia includes both numbers and graphs: "Descriptive statistics provides simple summaries about the sample and about the observations that have been made. Such summaries may be either quantitative, i.e. summary statistics, or visual, i.e. simple-to-understand graphs." (http://en.wikipedia.org/wiki/Descriptive_statistics).

What we get from this is that we don't have to do statistics with numbers - visualizations are equally valid. This should make the people who are nervous about digital humanities feel more relaxed. (Of course, visualizations bring their own fears - after all, humanities always avoided diagrams, let alone graphs, and the fact that visualization is now allowed to enter humanities is already quite amazing. (For example, current Cambridge University Press auhor guide still says that illustrations can only be used if it’s really necessary, because they distract readers from following the arguments in the text.)

Galton_graph_1886
Francis Galton’s first correlation diagram, showing the relation between head circumference and height, 1886. Source: Michael Friendly and Daniel Danis, The Early Origins and Development of the Scatterplot. The article suggests that this diagram was an intermediate form between a table of numbers and a true graph.


However, we are not finished yet. Besides numbers and graphs, we can also summarize the data using parts of this data. In this scenario, there is no translation of one media type into another type (for example, text translated into numbers or into graphs). Although such summaries are not (or not yet) understood as belonging to statistics, they perfectly fit the definition of descriptive statistics.

For example, we can summarize a text with numbers such as the average sentence length, the proportion between nouns and verbs, and so on. We can also use graphs: for example, a bar chart that shows frequency of all words used in the text in ascending order. But we can also use some words from the text as its summary.

One example is the popular word cloud. It summarizes a text by showing us most frequently used words that are scaled in size according to how often they are used. It carries exactly the same information, as a graph which plots their frequencies - but if the latter foregrounds the graphic representation of the pattern, the former foregrounds the words themselves.

Another example is a phrase net technique available on manyeyes (http://www-958.ibm.com/software/data/cognos/manyeyes/page/Phrase_Net.html). It is a graph that shows most frequent pairs of words in a text. Here is a prase net which shows first 20 most frequent word pairs in Jane Austen's Pride and Prejudice (1813):


Austin

Interactive version of this graph which allows you to change parameters.


In the previous examples, the algoritms extracted some words or phrases from a text and organized them visually - therefore it possible to argue that they still belong to graph method of descriptive statistics. But consider now the technique of topic modeling which recently has been getting lots of attention in digital humanities (http://en.wikipedia.org/wiki/Topic_model; http://tedunderwood.com/2012/04/07/topic-modeling-made-just-simple-enough/). The topic model algorithm outputs a number of sets of semantically related words. Each set of words is assumed to represent one theme in the text. Here are examples of three such sets from the topic model of the articles in journal Critical Inquiry (source: Jonathan Goodwin, Two Critical Inquiry Topic Models, 11/14/2012).

meaning theory interpretation question philosophy language point claim philosophical sense truth fact argument knowledge intention metaphor text account speech

history historical narrative discourse account contemporary terms status context social ways relation discussion essay sense form representation specific position

public war time national city american education work social economic space people urban culture corporate building united market business


In my lab, we have been developing free software tools for summarizing large image and video collections. In some of our applications, the whole collection is translated into a single high-resolution visualization (so these are not summaries) but in others, a sample of an image set, or parts of the images are used as the summaries. For example, here is the visual summary of 4535 covers of Time magazine which uses one one pizel wide horizontal column from from each cover. It shows the evolution of the covers design from 1923 to 2008 (left to right), compressing 4535 covers into a single image:

YZ_Time_covers


(For a detailed analysis of this and other related techniques for what I call exploratory media analysis, and their difference from information visualization, see my article What is Visualization?)


We can also think of other examples of summarizing collections / artifacts in different media by using parts of these collections / artifacts. On many web sites, video is summarized by a series of keyframes. (In computer science, there is a whole field called video summarization devoted to the development of algorithms to represent a video by using selected frames, or other parts of a video.)

To summarize a complex image we can translate it into a monochrome version that only shows the key shapes. Such images have been commonly used in many human cultures.

picasso_selfport1907
Picasso. Self Portrait. 1907. A summary of the face which uses outlines of the key parts.

Some symbols can also act as a summaries. For example, modernity and industrialization were often summarized by images of planes, cars, gear, workers, and so on. Today in TV commercials, a network society is typically summarized by a visualization showing a globe with animated curve connecting many points.

mechanization_takes_command
Example of an object becoming a symbol. The gears stand in for industrialization. The cover of Siegfried Gedeon's Mechanization Takes Command (1947).


There is nothing "positivist" or "scientific" about such same-media summaries, because they are what humanities and the arts have always been about. Art images always summarized visible or imaginary reality by representing only some essential details (like contours) and omitting others. (Compare to the goal of descriptive statistics to come up with "the basic features of the data.") A novel may summarize everything that happened over twenty years in the life of characters by only showing us a few of the events. And every review of a feature film includes a short text summary of its narrative.

Until development of statistics in the 19th century, all kinds of summaries were produced manually. Statistics "industrializes" this process, substituting subjective summarization by the objective and standardized measures. While at first these were only summaries of numerical data, the development of computational linguistics, digital image processing, and GIS in the 20th century also automates production of summaries of media such as texts, images, and maps.

Given that production of summaries is the key characteristics of human culture, I think that such traditional summaries created manually should not be opposed to more recent algorithmically produced summaries such as a word cloud or a topic model, or the graphs and numerical summaries of descriptive statistics, or binary (i.e., only black and wite without any gray tones) summaries of photographs created with image processing (in Photoshop, use Image > Adjustments > Threshold). Instead, all of them can be situated on a single continuous dimension.

On the one end, we have the summaries that use same media as the original "data." They also use the same the same structure (e.g., a long narration in the film or a novel is summarized by a condensed narrative presented in a few sentences in a review; a visible scene is represented by the outlines of the objects). We can also put here metonymy, (the key rhetorical figure), and Pierce's icon (from his 1867 semiotic triad icons/index/symbol).

On the other end, we have summaries which can use numbers and/or graphs and which present information which is impossible to see immediately by simply reading / viewing / listening / interacting with the cultural text or a set of texts. Examples of such summaries are a number representing the average number of words per sentence in a text which has tens of thousands of sentences, or the graph showing relative frequencies of all the words appearing in a long text.

I don't think that we can find some hard definite threshold which will separate summaries which can only be produced by algorithms (because of the data size) from the ones produced by humans. In the 1970s, without the use of any computers, French film theorists advanced the idea that most classical Hollywood films follow a single narrative formula. This is just one example of countless "summaries" produced in all areas of humanities (whether they are accurate is a separate question).

I don't know if my arguments will help us when we are criticized by people who keep insisting on a wrong chain of substitutions: digital humanities=statistics=science=bad. But if we keep explaining that statistics is not only about inferences and numbers, gradually we will be misunderstood less often.


(Note: in the last fifteen years, new mathematical methods for analyzing data which overlap with statistics become widely used – referred by "umbrella" terms uch as “data mining," "data science" and “machine learning.” They have very different goals than either descriptive or inferential statistics. I will address their assumptions, their use in industry, and how in digital humanities we may use them differently in a future article).


Monday, November 12, 2012

Image now: it is not cinema, or animation, or visual effects

(The following text is adapted from my next book Software Takes Command, forthcoming from Bloomsbury Academic in 2013)


Psyop - panosonic anthem 2012 - montage 2x2
Panosonic anthem commercial by Psyop, 2012.


TV commercials, television and film titles, and many feature films produced since 2000 feature a highly stylized visual aesthetics supported by animation and compositing software. Many layers of live footage, 3D and 2D animated elements, particle effects, and other media elements are blended to create a seamless whole. This result has the crucial codes of realism (perspective foreshortening, atmospheric perspective, correct combination of lights and shadows), but at the same time it enhances visible reality. (I can’t call this aesthetics “hyperreal” since the hybrid images assembled from many layers and types of media look quite different from the works of the hyperrealist artists such as Denis Peterson that visually look like standard color photographs.) Strong gray scale contrast, high color saturation, tiny waves of particles emulating from moving objects, extreme close-ups of textured surfaces (water drops, food products, human skin, finishes of consumer electronics devices, etc.), the contrasts between the natural uneven textured surfaces and smooth 3D renderings and 2D gradients, the rapidly changing composition and camera position and direction, and other devices heighten our perception. (For good examples of all these strategies, you can, for example, look at the commercials made by Psyop.)

We can say that they create a “map” which is bigger than the territory being mapped, because they show you more details and texture spatially, and at the same time compress time, moving through information more rapidly. We can also make another comparison with the Earth observation satellites which circle the planet, capturing its whole surface in detail impossible for any human observer to see – just as a human being can’t simultaneously see the extreme close-up of the surfaces and details of the movements of objects presented in the fictional space of a commercial.

None of the 20th century terms we inherited to talk about moving images describes this aesthetics, or production processes involved in creating it. It is not cinema, animation, special effects, invisible effects, visual effects, or even motion graphics. And yet, it characterizes the image today, and calls for its theoretical analysis and appreciation.

It is not cinema because live action is only a part of the sequence, and also because this live action is overlaid with 2D and/or 3D elements, and additional imagery layers. It is not animation or motion graphics because live action is central, as opposed to being just one element in a sequence. It is not special effects, because every frame in a sequence is an "effect" (as opposed to only selected shots). It is not invisible effects because the artifice and manipulations are made visible. It is not visual effects defined as "various processes by which imagery is created and/or manipulated outside the context of a live action shoot" - because here live action is manipulated.

Rather than seeing this new aesthetics, and production processes involved in its creation, as an extension of special effects / visual effects model, I think that it is more appropriate to see as an extension of the job of cinematographer of cinematography. 20th century cinematographer was responsible for selecting and choreographing all material elements which together produced the shots seen by the audiences: film stock, camera, lenses, lens filters, depth of field, focus, lighting, camera movements. Today a shot is likely to include many other elements and processes - image processing, compositing, 2D and 3D animation and models, relighting, particle systems, camera tracking, matte creation, effects, etc. (See for example the lists of features in Autodesk Flame software). However, at the end result is similar to what we had in the 20th century: a 3D scene that includes (real or constructed in 3D) bodies and objects. While now it also incorporates all kinds of digital transformations, and layers, it is still defined by three-dimensionality, perspective and MO movement of objects, just as the first films by The Lumiere brothers.

Psyop - Fage Plain commercial - 2011 - montage
Fage Plain commercial by Psyop, 2011.


Thursday, September 27, 2012

Search 357,000 TV broadcasts on Internet Archive

Internet_Archive_TV_interface
Internet Archive TV dataset - search interface.


Search 357,000 TV broadcasts on Internet Archive - and use their built-in visualizations to understand results:

http://archive.org/details/tv

Internet Archive's new sleek interface allows you to search its massive collection of TV broadcasts across 21 stations, from CNN to TeleFutura. For each search topic, the interface shows the graph of frequency over time; for each video, you can also see a bar graph of frequency of extracted topics. For example:

http://archive.org/details/CNNW_20111222_060000_Anderson_Cooper_360#start/85.5/end/115.5

I have argued that web interfaces to large cultural data sets need to incorporate visualizations - so you can get an overview of contents before deciding what to focus on. Internet Archive new interface is certainly a move in the right direction which will hopefully stimulates others to follow.

Wednesday, September 19, 2012

How many people in the world use Photoshop?

Facebook_social_graph
Facebook social graph.


How many people in the world use Photoshop? What about Illustrator, Final Cut, Maya, or any other popular media authoring application? How would we make even most approximate estimate? While Adobe, Autodesk, and Apple know how many copies of their software products they sell every quarter, this information is never released. But even if it was, the proportion of people who buy their software vs. all other users is so small that the sales numbers would not help us much.

We know so little about contemporary culture. Its flows, evolutionary mechanisms, patterns of dissemination, reuse, copying, and material bases (such as the number of copies of top software applications which enable it) are not visible to us. It is as though we have our eyes stuck very close to a map - we see few points, but unable to zoom out and see the larger picture.

One area where numbers and maps do exist is the usage of social networks. We know how many blogs are active, how many people use Twitter, how many photos are uploaded to Facebook every month (this number is already over 10 billion); YouTube gives us graphs showing the number of views for every video over time; Google Insights for Search allows us to study the popularity of any search term across time (since 2004) and territories. This is something, but its still limited. It is as though we zoomed out but are only able to see the overall contours on the map defining the areas - but not the details inside. And when some details are made available, they follow the rule of majority - we are told which topics, search terms, videos which are "most popular." These are the tips of the tips of the iceberg, and they are not very interesting. (For example, entering "top searches in the U.S. over 30 days" into Google Insights for Search returns these top items: 1) shoes, 2) samsung, 2) amazon, 4) boots.)

Back to Photoshop. Since late 1990s, I mostly work in cafes - which in Southern California means Starbucks. I am always curious what other people are doing on their computers and tablets around me, and I noticed the following pattern. If I am in a Starbucks and there are at least 10 other people with computers, one of them is using Photoshop. This is an informal observation, and it may only hold for the particular part of San Diego where I leave. But even if the real number is more like 1:20, this is already quite amazing.

What about the places where you live? Did you notice any similar pattern? if we can compare observation, it will give us at least some indication of how many people around the world are engaged in "art" and "design." Would not you want to know this?

Saturday, May 5, 2012

The Evolution of Video Game Controllers visualization


SOURCE: visual.ly

Browse more Gaming infographics.



Sunday, April 8, 2012

visualizing explosion of digital data


The World's Technological Capacity to Store, Communicate, and Compute Information.

Martin Hilbert1 and Priscila López.

Science, February 10, 2011.


Abstract:

We estimate the world's technological capacity to store, communicate, and compute information, tracking 60 analog and digital technologies during the period from 1986 to 2007. In 2007, humankind was able to store 2.9 × 1020 optimally compressed bytes, communicate almost 2 × 1021 bytes, and carry out 6.4 × 1018 instructions per second on general-purpose computers. General-purpose computing capacity grew at an annual rate of 58%. The world's capacity for bidirectional telecommunication grew at 28% per year, closely followed by the increase in globally stored information (23%). Humankind's capacity for unidirectional information diffusion through broadcasting channels has experienced comparatively modest annual growth (6%). Telecommunication has been dominated by digital technologies since 1990 (99.9% in digital format in 2007), and the majority of our technological memory has been in digital format since the early 2000s (94% digital in 2007).



Illustration from the article in Washington Post about this research:

Rise-of-Digital-Information



Friday, March 16, 2012

Computational humanities vs. digital humanities

Bluefinlabs
Bluefin Labs analyze 5 billion of online comments and 2.5 million minutes of TV every month.
This visualization shows the relations between Gatorade brand and the male viewers of different TV shows.
Source: Bluefin Mines Social Media To Improve TV Analytics, Fast Company, 11-07-2011.



echnonest_platform
Echonest offer information on 30 million songs and 1.5 million music artists.
Source: the.echonest.com



Facebook_social_graph
Paul Butler's visualization of a sample of 10 million friends from Facebook, using company' data warehouse.
Source: Paul Butler, Visualizing Friendships.




In a article called Computational Social Science (Science, vol. 323, no. 6, February 2009, the leading researchers in network analysis, computational linguistics, social computing, and other fields which now work with large data write:

"The capacity to collect and analyze massive amounts of data has transformed such fields as
biology and physics. But the emergence of a data-driven 'computational social science' has been much slower. Leading journals in economics, sociology, and political science show little evidence of this field. But computational social science is occurring in Internet companies such as Google and Yahoo, and in government agencies such as the U.S. National Security Agency. Computational social science could become the exclusive domain of private companies and government agencies. Alternatively, there might emerge a privileged set of academic researchers presiding over private data from which they produce papers that cannot be critiqued or replicated. Neither scenario will serve the long-term public interest of accumulating, verifying, and disseminating knowledge."

Substitute the word humanities in the above paragraph, and it now describes perfectly the issues involved in large-scale analysis of cultural data. Today digital humanities scholars are mostly working with the archives of digitized historical cultural archives which were created by libraries and universities with the funding from NEH and other institutions. These archives and their analysis is very important - but this work does not engage with the massive amounts of cultural content and peoples' conversations and opinions about it which exist on social media platforms, personal and professional web sites, and elsewhere on the web. This data offers us unprecedented opportunities to undertand cultural processes and their dynamics and develop new concepts and models which can be also used to better understand the past. (In our lab, we refer to computational analysis of large contemporary cultural data as cultural analytics.)

Contemporary media and web industries are dependent on the analysis of this data. This analysis enables search, recommendations, video fingerprinting, identification of trending topics, and other crucial functions of their services. Because of its scale and technical sophistication, perhaps we should call it "computational humanities." The players in computational humanities are Google, Facebook, YouTube, Bluefin labs, Echonest, and other companies which analyze social media signals (blogs, Twitter, etc.) and the content of media on social networks. They do not usually ask theoretical questions which can be directly related to humanities, but the types of analysis they perform and the techniques they use can be easily extended to ask these questions.

The questions posed in the paragraph I quoted above are directly applicable to "computational humanities." We can ask: Will computational humanities remain the exclusive domain of private companies and government agencies? Will we see a privileged set of academic researchers presiding over private data from which they produce computational humanities papers that cannot be critiqued or replicated?

These questions are essential for the future of humanities. In this respect, NEH/NSF Digging Into Data competitions are very important as they try to push humanists to think on the scale of computational humanities, and collaborate with computer scientists. To quote from the description of 2011 competition:

"The idea behind the Digging into Data Challenge is to address how "big data" changes the research landscape for the humanities and social sciences. Now that we have massive databases of materials used by scholars in the humanities and social sciences -- ranging from digitized books, newspapers, and music to transactional data like web searches, sensor data or cell phone records -- what new, computationally-based research methods might we apply?"

In our lab, we are hoping to make a contribution towards bridging the gap between "digital humanities" and "computational humanities." Our data sets range from the small historical datasets - for instance, 7000 year-old stone arrow heads and paintings of Piet Mondrian and Mark Rothko - to large scale contemporary user-generated content such as 1,000,000 manga pages or 1,000,000 images from deviantArt (the most popular social network for user-generated art). We also write papers for both humanities and computer science audiences. All our work is collaborative, involving students in digital art, media art, and computer science. And although our largest image sets are still tiny in comparison to the data analyzed by the companies I mentioned above, they are much bigger than what humanists and social scientists usually work with. The new visualization tools we have developed already allow you to explore patterns across 1,000,000 images, and we are gradually scaling them up.



Thursday, March 15, 2012

The state of Wikipedia infographic

Jen Rhee's infographic - How Wikipedia redefines how people do research.

my favorite data point: %56 of students will halt research if little information is found on wikipedia.

Wikipedia
Via: Open-Site.org

Wednesday, March 7, 2012

more museums put their collections online


A screenshot of SFMOMA ArtScope interface to their image collection developed by Stamen Design.


When we started our lab in 2007, we expected that in a few years massive sets of images of artworks (and other areas of visual culture) will start become available from major cultural institutions. So we focused on developing methods and techniques for their analysis (what we call cultural analytics), while waiting for these collections to become available. The wait is almost over - these collections are here. (Unfortunately right now museum interfaces often tell web visitors how many "objects" they have in their database, but not how images, so some of the numbers below are approximate.)

Examples:


MoMA - currently 33593 images online
www.moma.org/explore/collection/


Brooklyn Museum - probably 97,000 images online (their interface says that they have that many records but does not explain how many images they have)
http://www.brooklynmuseum.org/opencollection/collections/


SFMOMA - all 6,038 works in the collection are online:
http://www.sfmoma.org/projects/artscope


BBC Your Painting - currently already 110,000 images online from UK national collections, with the aim to reach 200,000
http://www.bbc.co.uk/arts/yourpaintings/


Cleveland Museum of Art -
currently 65,000 images online.

Whitney Museum, NYC:
not clear how many images are online but looks like a lot


Powerhouse Museum, Sydney - somewhere between 98,000 and 170,000 images online (their interface does not explain how many images they have)
http://www.powerhousemuseum.com/collection/database/



More and more museums offer their data via APIs:

List of museums offering API


The first museum API was developed some time ago by Powerhouse Museum - currently it offers access to 90,000 thumbnails:
http://www.powerhousemuseum.com/collection/database/download.php

Tuesday, February 28, 2012

world diary: how many photos do we take now?



From the talks at Digital Media Analysis, Search and Management workshop at Calit2, 2/27-2/28, 2012

talk 1:
Facebook - 3 billion photo uploads a month
Flickr - 5 billion photos in total
Flickr - %85 of photos have no tags or descriptions
Picassa - only %5 of photos contain tags
%60 of consumer photos contain faces

talk 2:
500 billion consumer photos taken per year
36 billion Facebook photos now
40K photos uploaded per minute

Sunday, February 26, 2012

an experiment on 283 million Facebook users: computational social science in action


Bakshy, Marlow, Rosenn, Adamic. The Role of Social Networks in Information Diffusion.

"At the time of the experiment [8/4/2010 - 10/4/2010] there were approximately 500 million Facebook users logging in at least once a month. The experimental population consisted from approximately 283 of these users."





Information diffusion in social networks: related research articles on Google Scholar.

Wednesday, February 22, 2012

a list of Innovative visualizations of temporal processes


I put together this online resource containing particularly interesting visualizations of temporal processes and some tools for time visualization:


LIST of INNOVATE VISUALIZATIONS OF TEMPORAL PROCESSES




Three examples of innovative visualizations of temporal cultural processes
(Movie narrative chart, Listening Post, The Ebb and Flow of Movies):

time_visualizations_examples






Sunday, February 12, 2012

New York Times: The Age of Big Data

Yet another article on big data - this one in the Sunday edition of New York Times:

Steve Lohr. The Age of Big Data. New York Times, February 12, 2012.


My favorite quotes from the article:

"“It’s a revolution,” says Gary King, director of Harvard’s Institute for Quantitative Social Science. “We’re really just getting under way. But the march of quantification, made possible by enormous new sources of data, will sweep through academia, business and government. There is no area that is going to be untouched.”

"Data is in the driver’s seat. It’s there, it’s useful and it’s valuable, even hip."











Definition of a data scientist

Do you enjoy finding patterns in big data?
Do you want the opportunity to apply the latest AI algorithms to real data?
Do beautiful data visualizations get you excited?
Do you want to have millions of people use your code every day?
Do you consider yourself an innovator? Tinkerer? Dreamer?

(Quoted rom the job description Data Scientist - Trulia Data Science Lab)

Saturday, February 11, 2012

Wolfram|Alpha Pro: Launching a Democratization of Data Science

At Software Studies Initiative, we have been working to democratize the use of digital image processing for exploring large image and video collections.

In September 20111, we released ImagePlot - a set of free open-source software tools which we developed in the lab and used in all our own research projects. You can now take a set of images and videos, automatically extract basic visual feature and then explore the patterns in your image or video collection using the extracted data.

So we are very exited to read the latest post from founder of Mathematica and Wolfram|Alpha Stephen Wolfram about the this week release of Wolfram|Alpha Pro, and how it automates analysis and visualization of the data sets.

http://blog.stephenwolfram.com/2012/02/launching-a-democratization-of-data-science/



Last October I was fortunate to chat with Stephen over lunch at Wolfram Data Summmit 2011 conference. He was interested in our work on analyzing and visualizing the spaces of variations of cultural artifacts. We talked about how the models of biological evolution and variability may related to cultural evolution and also the principles described in Wolfram's famous book The New Kind of Science.


Here are a few examples of our visualizations of spaces of variation in different kinds of cultural artifacts:

Mondrian vs Rothko: footprints and evolution in style space

Google logo space

One million manga pages

Friday, February 10, 2012

2012: social analytics for the rest of us


Until recently, large-scale social media analytics and large-scale user testing and was only done by large companies.

In 2012, expect many of these capabilities to become available to individual users - for relatively low rates (which means soon it will be free.)

Google pioneered with this Google Trends, Google Insights for Search, and Google Analytics (www.google.com/analytics) which since 2006 offered free in-depth analytics on an individual web site or blog.

According to one analysis, Google Analytics is now used on %50 of top 1,000,000 web sites. (Google Analytics Market Share. 2010-08-21.)

Google Trends and Insights were unique in allowing all of to see certain kinds trends as expressed by massive amounts of social data. In contrast, most other current offerings only allow you to analyze trends related to your own social product (web site, blog, Facebook page, YouTube video) performance.

For an example, see Facebook insights which "provides Facebook Page owners and Facebook Platform developers with metrics around their content."

And here is an example of user testing analytics - again, for your design:

Solidify (solidifyapp.com) is offering a few applications for web designers Verify is "the fastest way to collect and analyze user feedback on screens or mockups. See where people click, what they remember, or how they feel." Solidify is "the quickest way to prototype interface screens for user testing feedback. Learn where people get confused by page interactions."


This trend is not going to change overnight. However, more tools are coming and at least some of them promise free or almost free social analytics based on everybody's data:


Twitter "will unveil a series of new tools in the next few months, including sophisticated analytical tools, according to Erica Anderson, Twitter's manager for news and journalism." "Going forward, Anderson expects to see more people using Twitter to predict behaviors by analyzing wide swaths of tweets. She did not make it clear if the new analytical tools will include some features aimed at spotting trends. "The predictive nature of Twitter is still largely untapped," she said." (source: ReadWriteWeb, January 29, 2012.)

Socialflow (socialflow.com)is "Optimized Publisher for Twitter and Facebook. Publish items at the precise moments when they will maximize clicks, Retweets, mentions and follower growth. Use the SocialFlow AttentionScore™ to derive more value from your content."


Also, check out this experts discussion business-oriented trends in social analytics offerings.