opensource.google.com

Menu

RE2: a principled approach to regular expression matching

Thursday, March 11, 2010

Regular expressions are one of computer science's shining examples of the benefits of good computer science theory. They were originally developed by theorists as a way to describe infinite sets, but Ken Thompson introduced them to programmers as a way to describe text patterns in his implementation of the text editor QED for CTSS. Dennis Ritchie followed suit in his own implementation of QED, for GE-TSS. Thompson and Ritchie would go on to create Unix, and they brought regular expressions with them. By the late 1970s, regular expressions were a key feature of the Unix landscape, in tools such as ed, sed, grep, egrep, awk, and lex. They remain a key feature of the open source landscape today, in those venerable Unix tools and at the core of new languages like Perl, Python, and JavaScript.

The feature-rich regular expression implementations of today are based on a backtracking search with a potential for exponential run time and unbounded stack usage. At Google, we use regular expressions as part of the interface to many external and internal systems, including Code Search, Sawzall, and Bigtable. Those systems process large amounts of data; exponential run time would be a serious problem. On a more practical note, these are multithreaded C++ programs with fixed-size stacks: the unbounded stack usage in typical regular expression implementations leads to stack overflows and server crashes. To solve both problems, we've built a new regular expression engine, called RE2, which is based on automata theory and guarantees that searches complete in linear time with respect to the size of the input and in a fixed amount of stack space.

Today, we released RE2 as an open source project. It's a mostly drop-in replacement for PCRE's C++ bindings
and is available under a BSD-style license. See the RE2 project page for details.

Google Summer of Code: Applications Now Open for Mentoring Organizations

Monday, March 8, 2010


Looking for new contributors and fresh perspectives for your open source software project? Through the Google Summer of Code™ program, we fund students worldwide to work with mentors from the FLOSS community on a three month coding project. Over the past five years, we've successfully paired nearly 3,400 students "with more than 3,000 mentors from backgrounds spanning industry to academia, with some spectacular results: more than 8 million lines of source code produced and over $20M in funding in support of open source development. We're particularly excited by the social ties our students form through the course of the program. We've connected people in more than 100 countries, and hope to bring people from even more places into the Google Summer of Code community this year. We're looking forward to our sixth year and welcoming another group of 1,000 student developers to the program.

We're now accepting applications from open source projects who wish to act as mentoring organizations. We'll be taking mentoring organization applications until Friday, March 12th at 23:00 UTC. Our list of approved organizations will be published on the 2010 Google Summer of Code site on March 18th. Interested students will then have several days to discuss their ideas with the accepted organizations before student applications open on March 29th.

Check out our Frequently Asked Questions page for more details and a preview of the application. And remember, if you have any questions, you can always find us in the Google Summer of Code Discussion group or in #gsoc on Freenode. Best of luck to all of our applicants!

Make Contact with Google at SIGCSE 2010

Friday, March 5, 2010

Next week several Googlers will be attending and presenting at the 41st ACM Technical Symposium on Computer Science Education (SIGCSE 2010). From March 10-13, Leslie Hawthorn and Cat Allman from the Open Source Programs Office will be in Milwaukee, WI, USA to talk about Google’s open source student programs, Google Summer of Code™ and the Google Highly Open Participation Contest. Check out Google’s vendor session on Friday to hear more from Leslie and Cat. Leslie will also be speaking at a roundtable and panel discussion with Hal Abelson from the Google App Inventor team at the Humanitarian FOSS Symposium on Wednesday.

If you are interested in learning more about Google’s activities in computer science education, make sure to attend some of the talks we have scheduled or drop by the Google booth!

Low-Impact Operating System Tracing

The Google Open Source Team has the privilege of funding some really great projects in the Open Source space. Mathieu Desnoyers, a student at Ecole Polytechnique, recently defended his Ph.D. thesis, which we helped to fund. The topic of his thesis was "Low-Impact Operating System Tracing."

The open source projects he created as part of his work were two-fold: Linux Trace Toolkit Next Generation (LTTng), a LGPLv2.1/GPLv2 tracer for the Linux kernel; and Userspace RCU library (liburcu), a highly-scalable user-space synchronization library, distributed under the LGPLv2.1 license.

Mathieu was kind enough to send us this summary of his research:


Computer systems, both at the hardware and software-levels, are becoming increasingly complex. Tracing is the key to solving some or all of this increasing complexity. In the case of Linux, used in a large range of applications, from small embedded devices to high-end servers, the size of the operating system kernels are increasing, libraries are being added, and major redesign of existing software is required to benefit from multi-core architectures. As a result, the software development industry and individual developers are facing problems whose resolution requires an understanding of the interaction between applications and all components of an operating system.

In my thesis, I propose the LTTng (Linux Trace Toolkit next generation) tracer as an answer to the industry and open source community tracing needs. The low-intrusiveness of the tracer is a key aspect of its usefulness because we need to be able to reproduce problems occurring in normal conditions. In some cases, users leave tracers active at all times in production, which makes the tracer overhead definitely critical. Our approach involves the design of synchronization primitives that meet the low-impact requirements. The linearly scalable and wait-free RCU (Read-Copy Update) synchronization mechanism used by the LTTng tracer fulfills these requirements with respect to data read. A custom-made buffer synchronization scheme is proposed to extract tracing data while preserving linear scalability and wait-free characteristics.

By measuring the LTTng impact, I demonstrate that it is possible to create a tracer that satisfy all the following characteristics: low latency, deterministic real-time impact (wait-free), small impact on operating system throughput and linear scalability with the number of cores. Experiments on various architectures show that this tracer is portable.

I propose a general model for superscalar multi-core systems with weakly-ordered memory accesses to perform formal verification of the RCU correctness and wait-free guarantees by model-checking. The LTTng
buffering scheme is also formally verified for safety and progress. Formal verification demonstrates that these algorithms allow reentrancy from multiple execution contexts, ranging from standard thread to non-maskable interrupts handlers, allowing a wide instrumentation coverage of the operating system.


Many thanks to Mathieu for sending us this report. You can download the full dissertation for more details.

Google and the Tor Project

Thursday, March 4, 2010

When it comes to code, Google's support has made a big difference to the Tor Project. Providing privacy and helping to circumvent censorship online is a challenge that keeps our software developers and volunteers very busy. The Google Summer of Code™ brings students and mentors in the open source community together to write code for three months every year. A lot of coding got done in a few months in 2009, and Tor was lucky to get a group of students who kept on working past the summer months to improve existing projects and support users. Tor also works on Libevent with Google.

All of these changes in software are very exciting, but who is it all for? Why is anonymity online so important? Companies like Google have privacy and opt-out policies, but not everyone has this stance. Corporations, nations, criminal organizations and individuals want your information. Companies collect information on your web browsing habits and sell it or are sloppy when it comes to protecting it from identity thieves. Others can threaten lives, from repressive nations tracking down outspoken journalists, to abusive spouses or stalkers who want to find out where their victims are hiding; from enemy military forces trying to find a communications link, to criminals who know when law enforcement is watching online.



Political upheaval sparks protests and renewed efforts to control the flow of information online. Interest in censorship circumvention also rises. In 2009, use of Tor increased, as users tried to get around national firewalls during the elections in Iran, and after the introduction of national Internet filters in other countries.
In times of relative political stability, governments routinely filter out international news outlets, information on reproductive health, religion, human rights and other topics deemed unfit. Women blogging about things considered mundane elsewhere, like being forbidden to drive or shop alone, are harassed by authorities. On the one hand, technology has made it easier to crack down on dissent, but the right technology can influence policy in good ways. In Mauritania, the use of censorship circumvention software after 2005 became widespread enough to prompt the government to stop filtering, since it was becoming a waste of time.

Even people living in countries where free speech is protected by law need anonymity for political activities. People blogging about political views that differ from the prevailing attitudes in a small community may lose a job or face boycotts if they run a business. In a company town, writing about the misdeeds of the company that employs your neighbors may be dangerous. Telling people about corruption could lead to harassment from guilty officials.

When someone finds the courage to leave an abusive relationship, the support of victims' advocates is vital. The Internet can help a survivor find counseling, shelter, and encouragement from people who have gone through the same process. Sadly, stalkers are also using technology to find their victims. Abusers monitor web browsers to see if a victim is planning to leave. Information about a shelter's location can be found in email headers, forcing abuse survivors to relocate. According to the U.S. Bureau of Justice Statistics, over one in four people who are stalked experience some sort of cyberstalking. Though some software in a stalker's toolkit is installed on a home computer, IP addresses can reveal which internet cafe or library someone uses to get online. Even if you don't have a stalker, hiding your IP address can be a good idea. Kids and adults alike are advised not to tell strangers where they live, but an IP address can reveal it for them.

Sting operations fail if criminals can tell that the police are connecting to message boards and chat from a government network. The information disappears. Insurgents may be looking for soldiers connecting to their defense department's computers back home. Anonymous tip lines are not so anonymous if someone telling authorities about crime is the only person in the neighborhood connecting to a government website. Without anonymity, going after organized crime can be dangerous to officers and their families.

Some companies do not reveal how much they know about their customers, or who sees the information. Some Internet Service Providers feel entitled to sell data collected from their subscribers to marketers. Though they claim that the information is not tied to any particular users, it is easy to find someone based on their search history. Information about visits to banking websites, searches for details on pre-existing health conditions, or other sensitive online activity could be damaging in the wrong hands; whether made available through carelessness or commercial interest.

Privacy online can protect people offline whether they are organizing protests, covering the news, blowing the whistle on threats to public health, or just blogging about daily life. In the "real world" assaults on privacy like peeking in windows, opening mail, or breaking and entering are obvious crimes. In the online world, however, assaults on privacy are subtle and unyielding. These threats to your health, your wealth and your well-being have no "opt-out" button. They have no "scrub my data" option. Your online activities, e-mails, bank transactions and everything else can be used to trace where you are and who you are. Using software like Tor gives ordinary citizens more choice about the information they reveal online.

For more information about online privacy and circumventing internet censorship, visit the Tor Project's website.

.