Internship Description

Anunay Sanganal, ’20SEAS, interned with Zootera, a patent-pending platform that offers a direct, transparent and fun way to help protect wildlife and natural habitats. Anunay worked as a product engineer, creating a prototype for a business intelligence sustainability offering by implementing various techniques for data gathering, such as web scraping, database setup, and natural language processing techniques.

This summer, I interned at Zooterra as a product engineer. Zooterra is a startup in the environment conservation space and its mission is “to organize the world’s sustainability data to drive sustainable organizations and in order to protect our planet”. My job was to work on building a prototype business intelligence sustainability platform. According to research, there is a huge gap of $100B+ in order to protect forests and biodiversity around the world, and we need to act fast. Fortunately, impact investing continues to grow and a fourth of the $88 trillion professionally-managed assets are now allocated considering the ESG investment framework. However, while the financial and business communities are rapidly moving towards the ESG framework and funding sustainability and nature conservation opportunities, these opportunities are fragmented across the world and hard to access. The business intelligence platform will help simplify these fragmented pieces and make everything accessible in one place, allowing companies to prioritize nature conservation opportunities to drive more funding towards protecting the environment. My work was to gather information and data required for the companies tab on the platform, which would then be hosted on the Zooterra’s website. This data had two components: 1) sustainability information related to companies; and 2) news article mentions related to these companies. My work helped in developing and refining the platform that allowed the company to pitch the product to potential investors to raise seed funding.

For me, the internship was a great learning experience both in terms of technical and soft skills. The courses I had taken at Columbia during my fall and spring semesters had a huge impact on my internship. The work involved data aggregation from multiple sources from web, data manipulation, data analysis, machine learning, and natural language processing. My knowledge of web scraping, cloud computing, data manipulation, data analysis, and data visualization learned from data analysis, machine learning, and analytics on the cloud courses helped me tremendously in various aspects of the project. Further, as the data I was working on mostly involved text data, my knowledge from the course on natural language processing was of utmost help. Concepts like text cleaning and normalization, string matching (Fuzzy Matching), and named entity recognition and attention models (BERT) were used for extracting mentioned companies from new articles and information extraction from sustainability reports. Apart from the technical knowledge, the soft skills like communication, teamwork, organizing, and coordinating work with others, etc., developed by working on various course projects at Columbia helped me a lot for this internship. One of the most important soft skills that I learned was to communicate me ideas to a non-technical audience, a skill that I lacked.

I faced a lot of technical challenges during the course of the internship, as a lot of the work for building the platform we were trying to do was new to me, and there were very few resources online. One of the most significant roadblocks was extracting information from company sustainability PDF reports. Most of the existing packages in Python to extract text from PDF documents were not working as per our expectations. So, I built a pipeline from scratch that would take in a PDF and extract relevant text data using Python, OpenCV, Optical Character Recognition, and clustering algorithms (to determine the optimum parameters for OpenCV). Another roadblock we faced was of processes taking a lot of time and compute power. For this, I made use of my learnings fromthe Analytics on Cloud course to utilize Google Cloud Platform and multiprocessing and multitreading to speed up processes and save time. Another important challenge was presenting your results, technical issues and roadblocks, and plans/progress to a non-technical audience. This is an important skill as a data scientist, and I further improved it during the course of internship.

Overall, the internship was a great learning experience, both professionally and personally. The platform went live for pilot use for potential customers after two weeks of my internship. It will now be used to generate seed funding from potential investors and stakeholders. Seeing an idea that I worked on for two months come to fruition was an amazing feeling. I hope this platform is a success and that Zooterra will be able to raise funds for future expansion. I faced many roadblocks during my internship, but I was able to overcome them due to my learnings and knowledge gained from various courses at Columbia.

My key takeaways from this internship were the following:
- I improved my technical skills and worked on concepts that I had never worked on before.
- I learned how to effectively face challenges and roadblocks in projects. The best way to approach these challenges is to break them down into parts and address them one by one.
- I managed and organized my work, collaborating and working as a team and communicating with a non-technical audience.

I would like to thank Mr. Julio Corredor from Zooterra for giving me this wonderful opportunity and the Tamer Center for Social Enterprise for supporting me for this internship. It was an experience that will serve me well for the rest of my career!