Technophilic Magazine » Christopher Mitchell The voice of science and technology Wed, 07 Oct 2015 13:00:36 +0000 en-US hourly 1 http://wordpress.org/?v=3.8 Building a Virtual City from the Real World /2014/01/29/building-virtual-city-real-world/ /2014/01/29/building-virtual-city-real-world/#comments Wed, 29 Jan 2014 13:00:19 +0000 /?p=2001 From the early days of modern computing, the ability to simulate massive virtual worlds has been an attractive and lucrative concept. The games SimCity and Grand Theft Auto, featuring increasingly-elaborate worlds in each version, have sold millions of copies each.

The 1995 movie “Hackers” imagined a filesystem that looked like an intricate virtual city. Massively-multiplayer online games (MMOs) like World of Warcraft and multiplayer sandbox games like Minecraft have given players huge, even infinite worlds to explore or shape. In each case, the world is either created through tireless work by designers and graphic artists, or randomly generated by a procedural algorithm. The effort to build a virtual world that mimics real-world locations has invariably carried a prohibitive cost.

Three trends in modern computing have now come together to make detailed virtual copies of real-world locations possible. First, in an effort to support more detailed real-world navigation, many companies and academic groups have either manually created 3D models of buildings in large cities, or developed the technology to create such models automatically. Google’s Google Earth product contains a combination of automatically-generated buildings and models hand-created by users and employees. Microsoft and Here.com both offer mapping products with 3D buildings. The Computer Graphics and Immersive Technologies Laboratory at the University of Southern California has developed a method of creating 3D building models from LiDAR data. Second, extensive high-resolution orthoimagery (satellite imagery) and terrain information is freely available from sources such as the United States Geological Survey’s EROS service. Third, in the quest to satisfy the needs of so-called “Big Data”, a plethora of techniques have been developed to process large quantities of data in a highly parallel manner, vital for generating detailed virtual worlds in reasonable time.

fig1

On the hypothesis that data and technology has developed sufficiently to make the creation of virtual worlds from data about the physical world feasible, I have created a system called SparseWorld. SparseWorld combines orthoimagery, bathyspheric and elevation data from the USGS EROS service, and 3D buildings from Google’s 3D Warehouse. It can generate a full photorealistic terrain model of New York City with select buildings for the creative sandbox game Minecraft in a few hours on a cluster of servers containing an aggregate 300 cores and 200GB of RAM. In the remainder of this article, I will explain how SparseWorld works, discuss some of the most significant technical hurdles I faced, and conclude with a look forward on what’s next for SparseWorld.

How SparseWorld Works

As a proof-of-concept prototype, the SparseWorld system is constructed in Python, and combines components from several existing projects. The TopoMC project, itself built on a Python library called PyMCLevel, can generated scaled-down wilderness terrain from USGS elevation and groundcover data. TopoMC provided a base on which to code a converter that could also pull in satellite imagery, could generate full-scale terrain (in which one virtual meter equals one real-world meter), and could parallelize the terrain conversion across many cores or many machines. Because sample 3D building models from Google’s 3D Warehouse are available in the Collada format, I used the PyCollada library to parse the structure of 3D buildings and convert them into voxelized models appropriate for inclusion in a virtual world. I designed and implemented my own parallelization system on top of Python’s multithreading capabilities to allow terrain segments and buildings to be converted in parallel.

SparseWorld collects, converts, and combines the map and building datasets in six steps, five of which are currently implemented:

  1. GetRegion Phase: Determine what areas of real-world terrain data need to be fetched, download the relevant elevation, landcover, and orthoimagery data, and stitch together pieces as necessary.
  2. PrepRegion Phase: Warp and combine data into one large 8-layer GeoTIFF image. Layers are elevation, landcover, core depth, bathyspheric depth, terrain red channel, terrain green channel, terrain blue channel, and terrain IR channel.
  3. BuildRegion Phase 1: Generate Minecraft tiles (16 meter x 16 meter vertical slices of the terrain) from the terrain data (first half of BuildRegion phase). This phase is well-suited to parallelization.
  4. BuildRegion Phase 2: Weld tiles into regions, 512 meter x 512 meter vertical slices of the terrain, each of which is stored in a single file. This phase is reasonably well-suited to parallelization.
  5. StreetCorrect Phase: Generate 2D splines from OpenStreetMap data, correct building shadows and overlaps over streets in orthoimagery. (Planned)
  6. BuildingConvert Phase: Generate voxelized building, structure, and tree models from Collada 3D models, then place onto terrain. (Implemented, data missing) This phase will be parallelized with a pool of converter workers, a crossbar, and a pool of terrain region workers. (Planned)

In addition to these phases, it is clear that missing detail or errors in existing datasets will require some manual verification and tuning of the final results. While the architecture of the SparseWorld system was straightforward to design, the technical hurdles were significant, ranging from opaque file formats to game engine limitations to difficulty obtaining necessary datasets.

Technical Hurdles

Developing the SparseWorld system pushed me to solve several interesting subproblems, some related to my specialty of distributed computing, some in fields unfamiliar to me. The earliest challenge I ran across was converting mesh building models into properly-shaped and -textured voxel models. This was further complicated by details of the Collada model format that initially eluded me. Later, after I implemented terrain creation, conversions between different types of geographical coordinates led to small, hard-to-trace errors in terrain orientation and the alignment of buildings onto the terrain. Finally, while not strictly novel, developing an efficient method of parallelizing the terrain and building conversion required carefully considering the dependencies between the various phases of creating a world. I’ll focus on the process of voxelizing a mesh as the most interesting of the issues faced.

The process used to voxelize mesh models of buildings into voxel models appropriate for Minecraft went through several revisions and refinements. The very first version, a “strawman” or naïve implementation, computed the rectangular prism enclosing each mesh triangle, and filled any voxel that intersected with the prism with stone. This produced blocky but recognizable models, showing that I was correctly interpreting the original models’ mesh coordinates. The second refinement iterated through all of the voxels in the rectangular prism, only filling a voxel if the mesh triangle intersected that voxel. While this made most meshes convert properly, especially thin triangles would occasionally leave holes in the resulting model. A colleague specializing in scientific computing helped me develop a final refinement that iterates over the surface of the mesh rather than the volume of the enclosing prism, reducing the complexity of mesh conversion from O(n^3) to O(n^2) while improving accuracy and making texture-mapping much simpler.

What’s Next for SparseWorld

The current SparseWorld system is capable of generating 277 million square meters of terrain representing Manhattan Island and surrounding areas, comprising 71 billion cubic meters of compressed world information, in a few hours using a handful of powerful machines. The most pressing technical issue is the parallelization of the building conversion, which is at the top of the feature-implementation list. The biggest non-technical issue is the acquisition of a full 3D building dataset. While I have reached out to Google, Bing, and Here.com, I have had little success acquiring the necessary data to create a complete map of the city with all buildings.

Beyond those two issues, future development will focus on increasing the fidelity of the generated worlds. For example, a colleague has experimented with using computer vision techniques to recognize repeated patterns in buildings, which could be applied to identify windows in building models’ textures. Google Earth’s dataset includes accurate placement and models of trees in many cities, which would improve SparseWorld’s guesses of where trees should go based on landcover data and the infrared channel of USGS orthoimagery. Finally, to make the system performant, a compiled language rather than Python would be preferable, at the considerable development cost of re-implementing libraries like PyMCLevel and PyCollada.

More Information

]]>
/2014/01/29/building-virtual-city-real-world/feed/ 0
Self-Teaching a Love of STEM /2012/08/27/self-teaching-a-love-of-stem/ /2012/08/27/self-teaching-a-love-of-stem/#comments Mon, 27 Aug 2012 09:47:23 +0000 http://beta.technophilicmag.com/?p=403 Christopher Mitchell is currently pursuing a Ph.D. in Computer Science at NYU. Within the community of calculator programmers, he is a renowned afficionado. He recently published a book for beginners about programming the TI-83/84 graphing calculators, which you can order at manning.com/mitchell.

At the age of five, I was obsessed with trains. One day, my mother took me to the New York Transit Museum, which was holding a workshop about electricity, complete with batteries, light bulbs, magnet wire, and compasses. From that day, I knew I wanted to study electrical engineering, and began teaching myself about circuits, gadgets, and later, programming. I learned several programming languages, designed and built gadgets and hardware modifications, and eventually earned three degrees in electrical engineering and computer science. Relatively early in my programming and engineering career, I started to do teaching of my own, first online, and later for continuing education and undergraduate classes. I developed a conviction that all students and fledgling coders and engineers deserve the same opportunities to explore and teach on their own that I was given.

After my introduction to electronics at the Transit Museum, I sought out electronic design books, asked for Radio Shack’s then-admirable Forrest Mims 130-in-One and 300-in-One kits, and taught myself about components, circuits, and even the rudiments of the underlying math. In early elementary school, I learned LOGO and toyed with QBASIC, but I continued to mostly focus on designing and building circuits. Once I received my first graphing calculator, however, my focus began to shift.

When I was in seventh grade, I got a trusty TI-83 for Christmas. At first I thought it was little more than a fancy math tool, but as I began to use it more, I discovered it was something else entirely. I found that it had a program editor, and that by putting together a few commands, I could make the calculator do my bidding. The concepts of programming were not entirely foreign to me, thanks to my earlier exploration of LOGO and QBASIC, but I began to build a much greater breadth and depth of knowledge as I worked with my calculator. I started with simple animations, creating programs that drew and erased characters to create primitive ASCII art. After seeing a few games on friends’ calculators, I took faltering steps into what I later learned was reverse-engineering. I examined other programs, figured out how they worked, and used my new knowledge to improve and expand my own projects. I got involved in the international TI calculator programming and hobbyist community and published some of my projects. I fielded feedback, compliments, and criticism, and learned to grow as a person, a programmer, and even a marketer of my own work.

The TI-BASIC language that I had learned was easy but powerful, and I learned to do a great deal with it. However, I was frustrated to notice that some of the programs I encountered seemed far more advanced and powerful than anything I could make. When I tried to view their source code with the calculator’s built-in editor, I was confronted by a sea of random symbols. I eventually learned that these programs were written in z80 assembly language, created on a computer and assembled into a form the calculator could understand. Over a summer, I began to work with the language, at first painstakingly typing out the hexadecimal for each opcode on my calculator, and later gaining access to a computer to use an assembler.

By gradually honing my skills, I learned about the internals of processors, memory, and I/O, skills that matured into a love of low-level programming and hardware design as an undergraduate and graduate electrical engineer. I wrote a graphical shell for the TI-83+/84+ calculator, a mouse-based GUI library, games, a music and video player, and even a decentralized networking protocol. Looking back at all of my experiences with graphing calculator programming, I realize that it reinforced my enjoyment of working with circuits and taught me to enjoy hacking in the positive sense. I enjoy the challenge of making an extremely low-resource device do as much as possible, and pushing myself to complete projects that others might dismiss as impossible. Having taken myself from the simplest commands in BASIC to complex hand-coded z80 and x86 assembly, I decided I wanted to share my love of coding (and particularly calculator coding) with the masses.

When I was still taking my early steps with TI-BASIC, I founded an online forum and community website called Cemetech (“KE-me-tek”). I used it to publish my own programs and projects, but also began to use it to amass skilled and beginner programmers alike, who learned from each other and began to post their own projects. To date, Cemetech has amassed about three thousand users, and incubated software and hardware projects for calculators, computers, embedded systems, and the web. I was invited to teach beginner and advanced Java programming courses for my alma mater’s continuing education program. As a graduate student, I have twice taught my advisor’s undergraduate students about operating systems, C and x86 assembly programming, and reverse engineering. Their Computer Systems Organization class challenges them to launch buffer-overflow exploits, implement and optimize their own malloc design, and reverse-engineer raw x86, among other labs. I enjoy the challenge of teaching them these low-level concepts, and feel that the sense of accomplishment I feel when a student finally has a moment of understanding makes it worthwhile. I was therefore thrilled to be asked by Manning Publications to write a book about programming graphing calculators. “Programming the TI-83+/84+” is due in print this September, and is written to instill in readers young and old the same love of programming that I developed.

I believe that self-education and an early exposure to engineering and programming is vital to prompting a life-long love for these fields. Although I believe I had a predisposition towards technical fields and hobbies, the opportunities I was given fueled early interest that matured as I aged. In particular, without the ability to program my calculator constantly, whether at lunch, at home, or (perhaps unfortunately) during class and while walking to school, I doubt I’d have the love for and intuition into programming that I now have. In chatting with many current and ex-calculator programmers, I have heard countless versions of my own story: the self-driven exploration of the calculator’s features, the thirst to learn what made programs and games tick, the love of surmounting a good challenge. I think the burden lies on museums, libraries, and even technology companies to be good citizens and make such opportunities available to children and teenagers.

A year and a half ago, I wrote an editorial criticizing Texas Instruments, who had taken a nearly Apple-esque position in locking down their new TI-Nspire graphing calculator. Native programming in assembly, C, or even TI-BASIC was impossible with their new calculator, a restriction aimed at placating teachers upset about students playing games in class. Only via third-party hacks could the device be unlocked, and with each new operating system (OS) release, TI squashed the existing unlock exploits. In my piece, I decried TI’s attitude with the Nspire as astonishingly short-sighted; it yielded a calculator that would not allow students to explore programming, a stark contrast to the TI-83+/84+ series. Texas Instruments vociferously promotes STEM (Science, Technology, Engineering, and Mathematics) education, but requiring “jail-breaking” to even write usable programs on the calculators showed exactly the opposite attitude. Thankfully, my editorial and other negative press forced them to partially reverse their decision, and they have now made the Lua language writeable on the calculator. Nevertheless, the Nspire continues to largely cater to a narrow view of the needs of teachers, rather than encompassing the equally-important needs of the students who buy and use Texas Instruments’ calculators.

Perhaps you have a similar story of how you got into STEM fields, where you pushed yourself to learn from existing programs, from taking things apart, and from books. Even if your knowledge of technology and engineering comes entirely from formal instruction in classes, I believe you can appreciate the value of earlier exposure to the fields in every form, from hackable, programmable gadgets to easy-to-access fora with free expert programming help. I encourage educational and commercial institutions large and small to press forward in educating younger generations and to give them the opportunity to make the same self-driven discoveries that many of us once made.

]]>
/2012/08/27/self-teaching-a-love-of-stem/feed/ 0