Showing posts sorted by relevance for query searle chinese room. Sort by date Show all posts
Showing posts sorted by relevance for query searle chinese room. Sort by date Show all posts

Saturday, March 24, 2007

The Mechanical Turk and Searle's Chinese Room

The Times has an article about Jeff Bezos' Mechanical Turk project, which lets machines outsource certain tasks to humans. (The orginal mechanical Turk was an 18th century hoax in which a hidden human operated a chess-playing automaton.) As Bezos describes,

“Normally, a human makes a request of a computer, and the computer does the computation of the task,” he said. “But artificial artificial intelligences like Mechanical Turk invert all that. The computer has a task that is easy for a human but extraordinarily hard for the computer. So instead of calling a computer service to perform the function, it calls a human.”

...The company opened Mechanical Turk as a public site in November 2005. Today, there are more than 100,000 “Turk Workers” in more than 100 countries who earn micropayments in exchange for completing a wide range of quick tasks called HITs, for human intelligence tasks, for various companies.

The Times writer Jason Pontin (who is also editor and publisher of MIT's Technology Review), gives Turk working a try, and finds it disorienting:

What is it like to be an individual component of these digital, collective minds?

To find out, I experimented. After registering at www.mturk.com, I was confronted with a table of HITs that I could perform, together with the price that I would be paid. I first accepted a job from ContentSpooling.net that asked me to write three titles for an article about annuities and their use in retirement planning. Then I viewed a series of images apparently captured from a vehicle moving through the gray suburbs of North London, and, at the request of Geospatial Vision, a division of the British technology company Oxford Metrics Group, identified objects like road signs and markings.

For all this, my Amazon account was credited the lordly sum of 12 cents. The entire experience lasted no more than 15 minutes, and from my point of view, as an occluded part of the hive-mind, it made no sense at all.

This is reminiscent of philospher John Searle's thought experiment called the Chinese Room, in which he posits a large team of humans implementing an algorithm that translates Chinese to English. Since each human performs only a small task (e.g., sorting acording to a rule set), none have any understanding of the overall process. Searle asks where, exactly, does the understanding of Chinese and English reside in this device? Searle considered his thought experiment as evidence against strong AI, whereas I just consider Searle to be confused. It's obvious that a Turk worker might be a small cog in some larger process that "understands" the world and processes information in a useful way. This depends not at all on what the little cog understands or does not understand.

Wednesday, May 23, 2018

Dominic Cummings on Fighting, Physics, and Learning from tight feedback loops

Another great post from Dom.

Once something has become widely understood, it is difficult to recreate or fully grasp the mindset that prevailed before. But I can attest to the fact that until the 1990s and the advent of MMA, even "experts" (like boxing coaches, karate and kung fu instructors, Navy SEALs) did not know how to fight -- they were deeply confused as to which techniques were most effective in unarmed combat.

Soon our ability to predict heritable outcomes using DNA alone (i.e., Genomic Prediction) will be well-established. Future generations will have difficulty understanding the mindset of people (even, scientists) today who deny that it is possible.

The same will be true of AGI... For example, see the well-known "Chinese Room" argument against AGI, advanced by Berkeley Philosopher John Searle (discussed before in The Mechanical Turk and Searle's Chinese Room). Searle's confusion as to where, exactly, the understanding resides inside a complex computation seems silly to us today given recent developments with deep neural nets and, e.g., machine translation (the very problem used in his thought experiment). Understanding doesn't exist in any sub-portion of the network, it is embodied in the network. (See also Thought vectors and the dimensionality of the space of concepts :-)
Effective action #4a: ‘Expertise’ from fighting and physics to economics, politics and government

Extreme sports: fast feedback = real expertise

In the 1980s and early 1990s, there was an interesting case study in how useful new knowledge jumped from a tiny isolated group to the general population with big effects on performance in a community. Expertise in Brazilian jiu-jitsu was taken from Brazil to southern California by the Gracie family. There were many sceptics but they vanished rapidly because the Gracies were empiricists. They issued ‘the Gracie challenge’.

All sorts of tough guys, trained in all sorts of ways, were invited to come to their garage/academy in Los Angeles to fight one of the Gracies or their trainees. Very quickly it became obvious that the Gracie training system was revolutionary and they were real experts because they always won. There was very fast and clear feedback on predictions. Gracie jiujitsu quickly jumped from an LA garage to TV. At the televised UFC 1 event in 1993 Royce Gracie defeated everyone and a multi-billion dollar business was born.

People could see how training in this new skill could transform performance. Unarmed combat changed across the world. Disciplines other than jiu jitsu have had to make a choice: either isolate themselves and not compete with jiu jitsu or learn from it. If interested watch the first twenty minutes of this documentary (via professor Steve Hsu, physicist, amateur jiu jitsu practitioner, and predictive genomics expert).

...

[[ On politics, a field in which Dom has few peers: ]]

... The faster the feedback cycle, the more likely you are to develop a qualitative improvement in speed that destroys an opponent’s decision-making cycle. If you can reorient yourself faster to the ever-changing environment than your opponent, then you operate inside their ‘OODA loop’ (Observe-Orient-Decide-Act) and the opponent’s performance can quickly degrade and collapse.

This lesson is vital in politics. You can read it in Sun Tzu and see it with Alexander the Great. Everybody can read such lessons and most people will nod along. But it is very hard to apply because most political/government organisations are programmed by their incentives to prioritise seniority, process and prestige over high performance and this slows and degrades decisions. Most organisations don’t do it. Further, political organisations tend to make too slowly those decisions that should be fast and too quickly those decisions that should be slow — they are simultaneously both too sluggish and too impetuous, which closes off favourable branching histories of the future.




See also Kosen Judo and the origins of MMA.


Choking out a Judo black belt in the tatami room at the Payne Whitney gymnasium at Yale. My favorite gi choke is Okuri eri jime.


Training in Hawaii at Relson Gracie's and Enson Inoue's schools. The shirt says Yale Brazilian Jiujitsu -- a club I founded. I was also the faculty advisor to the already existing Judo Club :-)

Sunday, May 12, 2013

Dennett and Intuition Pumps

At the bookstore today I spent some time looking at Intuition Pumps And Other Tools for Thinking, Daniel Dennett's new book. I highly recommend his Darwin's Dangerous Idea, discussed earlier here. I'm not a big fan of Dennett's work on free will and determinism (for my views, see this old post and also here), but we seem to share the same opinion of John Searle's Chinese Room.

For more Dennett, see this Stanford Humanities Center lecture (iTunes video).
NYTimes: ... The new book, largely adapted from previous writings, is also a lively primer on the radical answers Mr. Dennett has elaborated to the big questions in his nearly five decades in philosophy, delivered to a popular audience in books like “Consciousness Explained” (1991), “Darwin’s Dangerous Idea” (1995) and “Freedom Evolves.”

The mind? A collection of computerlike information processes, which happen to take place in carbon-based rather than silicon-based hardware.

The self? Simply a “center of narrative gravity,” a convenient fiction that allows us to integrate various neuronal data streams.

The elusive subjective conscious experience — the redness of red, the painfulness of pain — that philosophers call qualia? Sheer illusion.

Human beings, Mr. Dennett said, quoting a favorite pop philosopher, Dilbert, are “moist robots.”

“I’m a robot, and you’re a robot, but that doesn’t make us any less dignified or wonderful or lovable or responsible for our actions,” he said. “Why does our dignity depend on our being scientifically inexplicable?”

If he hadn’t grown up in an academic family, Mr. Dennett likes to say, he probably would’ve been an engineer. From his beginnings in the philosophical hothouses of early 1960s Harvard and Oxford, he had a feeling of being out of step joined by a precocious self-confidence.

As an undergraduate, he transferred from Wesleyan University to Harvard so he could study with the great logician W. V. O. Quine and explain to him why he was wrong. “Sheer sophomoric overconfidence,” Mr. Dennett recalled.

As a doctoral student at Oxford, then the center of the philosophical universe, he studied with the eminent natural-language philosopher Gilbert Ryle but increasingly found himself drawn to a more scientific view of the mind.

“I vividly recall sitting with my landlord’s son, a medical student, and asking him, ‘What is the brain made of?’ ” Mr. Dennett said. “He drew me a simple picture of a neuron, and pretty soon I was off to the races.”

In 1969, Mr. Dennett began keeping his “Philosophical Lexicon,” a dictionary of cheeky pseudo-terms playing on the names of mostly 20th-century philosophers, including himself. (“dennett: an artificial enzyme used to curdle the milk of human intentionality.”) Today, his impatience with the imaginary games philosophers play — “chmess” instead of chess, he calls it — and his preference for the company of scientists lead some to question if he’s still a philosopher at all.

“I’m still proud to call myself a philosopher, but I’m not their kind of philosopher, that’s for sure,” he said. The new book reflects Mr. Dennett’s unflagging love of the fight, including some harsh whacks at longtime nemeses like the paleontologist Stephen Jay Gould — accused of practicing a genus of dirty intellectual tricks Mr. Dennett calls “goulding” — that some early reviewers have already called out as unsporting. (Mr. Gould died in 2002.)

Mr. Dennett also devotes a long section to a rebuttal of the famous Chinese Room thought experiment, developed by 30 years ago by the philosopher John Searle, another old antagonist, as a riposte to Mr. Dennett’s claim that computers could fully mimic consciousness.

Clinging to the idea that the mind is more than just the brain, Mr. Dennett said, is “profoundly naïve and anti-scientific.”

Sunday, April 19, 2009

50 years of John Searle at Berkeley

To find a 90 minute podcast of this gathering, which is remarkable for the quality of the speeches given in honor of philosopher John Searle, search under "searle 50 berkeley" at iTunes U (or follow this link).

John Searle’s 50 Years at Berkeley—A Celebration

A celebration of John Searle’s 50 years of distinguished service to the UC Berkeley campus, with reflections by Tom Nagel, Barry Stroud, Robert Cole, Alex Pines, Peter Hanks, and Maya Kronfeld.

While I disagree strongly with Searle's most famous philosophical construct -- the so called Chinese room argument against strong AI (see also here) -- I've always found his writing and argumentation to be exceptionally clear, at least for a philosopher ;-)

See also Paul Graham against philosophy.

Wednesday, December 14, 2016

Thought vectors and the dimensionality of the space of concepts


This NYTimes Magazine article describes the implementation of a new deep neural net version of Google Translate. The previous version used statistical methods that had reached a plateau in effectiveness, due to limitations of short-range correlations in conditional probabilities. I've found the new version to be much better than the old one (this is quantified a bit in the article).

These are some of the relevant papers. Recent Google implementation, and new advances:
https://arxiv.org/abs/1609.08144https://arxiv.org/abs/1611.04558.

Le 2014, Baidu 2015, Lipton et al. review article 2015.

More deep learning.
NYTimes: ... There was, however, another option: just design, mass-produce and install in dispersed data centers a new kind of chip to make everything faster. These chips would be called T.P.U.s, or “tensor processing units,” ... “Normally,” Dean said, “special-purpose hardware is a bad idea. It usually works to speed up one thing. But because of the generality of neural networks, you can leverage this special-purpose hardware for a lot of other things.” [ Nvidia currently has the lead in GPUs used in neural network applications, but perhaps TPUs will become a sideline business for Google if their TensorFlow software becomes widely used ... ]

Just as the chip-design process was nearly complete, Le and two colleagues finally demonstrated that neural networks might be configured to handle the structure of language. He drew upon an idea, called “word embeddings,” that had been around for more than 10 years. When you summarize images, you can divine a picture of what each stage of the summary looks like — an edge, a circle, etc. When you summarize language in a similar way, you essentially produce multidimensional maps of the distances, based on common usage, between one word and every single other word in the language. The machine is not “analyzing” the data the way that we might, with linguistic rules that identify some of them as nouns and others as verbs. Instead, it is shifting and twisting and warping the words around in the map. In two dimensions, you cannot make this map useful. You want, for example, “cat” to be in the rough vicinity of “dog,” but you also want “cat” to be near “tail” and near “supercilious” and near “meme,” because you want to try to capture all of the different relationships — both strong and weak — that the word “cat” has to other words. It can be related to all these other words simultaneously only if it is related to each of them in a different dimension. You can’t easily make a 160,000-dimensional map, but it turns out you can represent a language pretty well in a mere thousand or so dimensions — in other words, a universe in which each word is designated by a list of a thousand numbers. Le gave me a good-natured hard time for my continual requests for a mental picture of these maps. “Gideon,” he would say, with the blunt regular demurral of Bartleby, “I do not generally like trying to visualize thousand-dimensional vectors in three-dimensional space.”

Still, certain dimensions in the space, it turned out, did seem to represent legible human categories, like gender or relative size. If you took the thousand numbers that meant “king” and literally just subtracted the thousand numbers that meant “queen,” you got the same numerical result as if you subtracted the numbers for “woman” from the numbers for “man.” And if you took the entire space of the English language and the entire space of French, you could, at least in theory, train a network to learn how to take a sentence in one space and propose an equivalent in the other. You just had to give it millions and millions of English sentences as inputs on one side and their desired French outputs on the other, and over time it would recognize the relevant patterns in words the way that an image classifier recognized the relevant patterns in pixels. You could then give it a sentence in English and ask it to predict the best French analogue.
That the conceptual vocabulary of human language (and hence, of the human mind) has dimensionality of order 1000 is kind of obvious*** if you are familiar with Chinese ideograms. (Ideogram = a written character symbolizing an idea or concept.) One can read the newspaper with mastery of roughly 2-3k characters. Of course, some minds operate in higher dimensions than others ;-)
The major difference between words and pixels, however, is that all of the pixels in an image are there at once, whereas words appear in a progression over time. You needed a way for the network to “hold in mind” the progression of a chronological sequence — the complete pathway from the first word to the last. In a period of about a week, in September 2014, three papers came out — one by Le and two others by academics in Canada and Germany — that at last provided all the theoretical tools necessary to do this sort of thing. That research allowed for open-ended projects like Brain’s Magenta, an investigation into how machines might generate art and music. It also cleared the way toward an instrumental task like machine translation. Hinton told me he thought at the time that this follow-up work would take at least five more years.
The entire article is worth reading (there's even a bit near the end which addresses Searle's Chinese Room confusion). However, the author underestimates the importance of machine translation. The "thought vector" structure of human language encodes the key primitives used in human intelligence. Efficient methods for working with these structures (e.g., for reading and learning from vast quantities of existing text) will greatly accelerate AGI.

*** Some further explanation, from the comments:
The average person has a vocabulary of perhaps 10-20k words. But if you eliminate redundancy (synonyms + see below) you are probably only left with a few thousand words. With these words one could express most concepts (e.g., those required for newspaper articles). Some ideas might require concatenations of multiple words: "cougar" = "big mountain cat" , etc.

But the ~1k figure gives you some idea of how many distinct "primitives" (= "big", "mountain", "cat") are found in human thinking. It's not the number of distinct concepts, but rather the rough number of primitives out of which we build everything else.

Of course, truly deep areas of science discover / invent new concepts which are almost new primitives (fundamental, but didn't exist before!), such as "entropy", "quantum field", "gauge boson", "black hole", "natural selection", "convex optimization", "spontaneous symmetry breaking", "phase transition" etc.
If we trained a deep net to translate sentences about Physics from Martian to English, we could (roughly) estimate the "conceptual depth" of the subject. We could even compare two different subjects, such as Physics versus Art History.

Wednesday, May 27, 2020

David Silver on AlphaGo and AlphaZero (AI podcast)



I particularly liked this interview with David Silver on AlphaGo and AlphaZero. I suggest starting around ~35m in if you have some familiarity with the subject. (I listened to this while running hill sprints and found at the end I had it set to 1.4x speed -- YMMV.)

At ~40m Silver discusses the misleading low-dimensional intuition that led many to fear (circa 1980s-1990s) that neural net optimization would be stymied by local minima. (See related discussion: Yann LeCun on Unsupervised Learning.)

At one point Silver notes that the expressiveness of deep nets was never in question (i.e., whether they could encode sufficiently complex high-dimensional functions). The main empirical question was really about efficiency of training -- once the local minima question is resolved what remains is more of an engineering issue than a theoretical one.

Silver gives some details of the match with Lee Sedol. He describes the "holes" in AlphGo's gameplay that would manifest in roughly 1 in 5 games. Silver had predicted before the match, correctly, that AlphaGo might lose one game this way! AlphaZero was partially invented as a way to eliminate these holes, although it was also motivated by the principled goal of de novo learning, without expert examples.

I've commented many times that even with complete access to the internals of AlphaGo, we (humans) still don't know how it plays Go. There is an irreducible complexity to a deep neural net (and to our brain) that resists comprehension even when all the specific details are known. In this case, the computer program (neural net) which plays Go can be written out explicitly, but it has millions of parameters.

Silver says he worked on AI Go for a decade before it finally reached superhuman performance. He notes that Go was of special interest to AI researchers because there was general agreement that a superhuman Go program would truly understand the game, would develop intuition for it. But now that the dust has settled we see that notions like understand and intuition are still hidden in (spread throughout?) the high dimensional space of the network... and perhaps always will be. (From a philosophical perspective this is related to Searle's Chinese Room and other confusions...)

As to whether AlphaGo has deep intuition for Go, whether it can play with creativity, Silver gives examples from the Lee Sedol match in which AlphaGo 1. upended textbook Go theory previously embraced by human experts (perhaps for centuries?), and 2. surprised the human champion by making an aggressive territorial incursion late in the game. In fact, human understanding of both Chess and Go strategy have been advanced considerably via AlphaZero (which performs at a superhuman level in both games).

See also this Manifold interview with John Schulman of OpenAI.

Blog Archive

Labels