Internship Description

Shangyou Wu, ‘19SEAS, was a data scientist intern at GrowSquares. He was a key player in the development of plant recommendation and support engines of GrowSquares’ digital products and services. Shangyou was responsible for defining the data structures GrowSquares used to determine which plants work best, and ensure they are performing when used by customers. He was also primarily responsible for improving the engine behind plant success, developing tests that isolate susceptibility to varying disease/fungi and improve plant yield.

Since I’m working for Growsquares who’s in its early stage, I’m responsible for all data science projects including building predictive models for sunlight and temperature, Finding the best way to water plants by establishing several algorithms and optimized the gardens layout based on compatibility. Besides, in order to finalize my work and help launch this product by the end of August, I’m also responsible for communicating with software developer and completed the interaction between algorithms and databases.

The recommendation system and monitoring algorithms I built are the core values of Growsquares. This system can help users figure it out which plants can grow well in their places and the monitoring algorithms can give precise instructions about how to grow and when to water. Since this company defines itself as “provider of smart personal gardens”, the recommendation system shows that GrowSquares understands users’ different situations and we’re doing some advanced calculations based on their own information in order to build the best customised gardens. As for the monitoring system, it can add more humanizing flavour to our company image because it shows that we actually take care of their plants after their purchases. We send notifications when the plants need watering, we ask them to check if the plants enter the next grow stage or not. All we do is to make sure our users have the best gardens they could ever have.

The recommendation system is the first thing our users will use and it’s also the foundation of all other data science projects. This recommendation system consists of two parts, Plants ranking for users and the perfect layout of users’ personal gardens.

For plants ranking, since different plant has different ideal environment to grow. We wanted to know that for the next two months, what will the weather look like and which plant can grow the best in that situation. We found out that temperature and sunlight intensity have the dominant impact on plants growth status, so our business goal comes to predict the temperature and sunlight intensity for the next month. For temperature, because we wanted to predict not only the average but also the maximum and minimum temperature, so traditional time series models cannot be applied in this problem. We decided to use deep learning model to learn the trend hiding behind the complicated historical data. Sunlight intensity prediction is another special case. I ended up with a tree-based model based on other weather features to predict the sunlight intensity in the future, the details are also be elaborated in the section about how does this internship contribute to my coursework.

Once we get the user’s environment information, we rank plant by how well they can grow under the situation. Based on the ranking score, we set the priority for different plants and help our users to figure it out how to arrange these plants. For example, some plants may be harmful to other plants so they should not be next to each other.

I encountered lots of problems when I tried to apply different models into our business problems, so I had to give an alternative solution while taking it into account operability. For example, when we wanted to predict the sunlight intensity for the next two months, the most ideal way was to apply time series models based on the historical data of sunlight intensity. However, we only have the historical data that ends in 2017, which means we don’t have the most recent data. Time series model can only be applied to consecutive dataset so I had to think of a new way to predict sunlight intensity. Then I found another dataset including the most recent weather data. Unfortunately, even though it did include some sunlight related features like cloud coverage and day time, it didn’t include the sunlight intensity. Since I can get the sunlight intensity information from the first dataset and the most recent other information from the second dataset, I came up with a way to predict the sunlight intensity based on features I can get from the second dataset. Firstly, I combined these two datasets using their overlapped parts, old historical data until 2017. Then I set the sunlight intensity as the dependent variable and other features from the second dataset as independent variables. In this way, our model can find the relationship between the features from the second dataset and sunlight intensity. So now I can easily predict sunlight intensity based on the model and the second dataset.

As an early stage startup, Growsquares gave me an opportunity to know the whole process of how to build a real thing from scratch. We need to conceptualize everything and come up with alternatives in case some accidents happen. I need to evaluate every decision teams make while considering realistic situation so I won’t waste too much time working on useless contents since our time is tight, team is small and we have to launch the product this summer. I also got to know how does each function in a company interact with others. I used to think meeting is not helpful for my work since I just need to know the inputs and outputs of my models. This internship totally changed my mindset and made me realize that communication is the key to improve working efficiency. When I was working on watering algorithms, the support from backend engineer and access to databases is necessary and we won’t make it if we don’t communicate.

This internship also helped me deepen my understanding of what kind of hard skills industries are looking for and what should I keep working on to find an ideal full-time job in the next semester.

Working for a startup is a precious experience for me to think about what I truly want and how I build my career path. This experience helped me expand my imagination and solidify my determination for my future career to be a venture capital investor focusing on artificial intelligence area. I really appreciate the company and everyone I met there.