Anna completed her veterinary training in 2012 and worked for a few years before returning to SLU in 2018 to begin her PhD studies at the Department of Animal Nutrition and Management (now the Department of Applied Animal Science and Welfare).
– I worked mainly on the INDILACT research project. The project is still ongoing, funded by the foundation Seydlitz MP bolagen, the Swedish farmers’ foundation for agricultural research and Formas. The aim is to improve the health and fertility of dairy cows whilst also improving milk production, by developing tools to tailor how often a cow calves based on the cow’s and the farm’s specific conditions.
In the project, Anna worked with large amounts of data and extensive datasets, which came from a variety of sources. Much of the data came from the consultancy firm Växa’s database Kokontrollen. Kokontrollen contains data from Swedish farms that is used as the basis for advising Swedish farmers.
– The database includes information on when the cows are inseminated, when they calve, how much milk they produce during test milkings, and the composition of the milk. Everything I was interested in!
How to work with data was something Anna felt was missing from her undergraduate studies.
– It might have come up in a lecture, with someone talking about the importance of openness and transparency in research. But not how to actually do it in practice.
As a PhD student, Anna took the library’s PhD course ‘Information retrieval and methods for scientific communication’.
– It was one of the best courses I took as a PhD student. I picked up lots of useful tools, including the basics of how to handle data. It was incredibly useful.
Anna also took the course ‘Advanced use of Excel’ and found it very useful. However, she wishes there had been more courses on data management, for example on how to organise data, create a folder structure and write metadata clearly.
– For me, data management was such a big part of my PhD that I would have liked an even more in-depth course. If I had started with a course like that, just imagine how much time I could have saved!
Anna has published articles primarily in the Journal of Dairy Science, which does not want data published as supplementary material but rather in a repository. That was one of the reasons why she published her data openly, apart from the fact that she believes it is important.
– Good data management is the foundation of good research. Otherwise, it is difficult to trust the results. I believe that research should be as transparent and as accessible as possible.
Anna also mentions the FAIR principles, which are international guidelines for making research data as usable and reusable as possible.
– I think the FAIR principles are extremely important. Sharing results and methods so that others don’t have to reinvent the wheel is a major part of research work. The fact that something can be used by others saves time and means that knowledge development progresses more quickly.
Anna published data from her projects in a repository managed by the Swedish National Data Service (a research infrastructure that SLU is part of). As she had not published data before, she sought help from the SLU University Library and SLU’s data management support service.
– The help I received was fantastic. With everything, every step of the way, calmly and professionally.
As much of the data Anna was working with came from a company, she drew up an agreement on how it could be used. For example, she had received data from 20 farms and, under the agreement, was not permitted to publish information that could be linked to an individual farm. Instead of making all the data available. Anna therefore published metadata, i.e. descriptions of the data.
– This database, Kokontrollen, is incredibly important, both for Växa and for the farms themselves. We at SLU are also keen to maintain a good working relationship with them. The data they collect is extremely valuable to us, and we naturally hope that our research is also important to them.
Anna’s experience of working with Kokontrollen also became a key part of the development of SLU’s cow data infrastructure, Gigacow. A collaboration that now continues in her new role at the Seydlitz Laboratory at SLU, where she works on utilising and making available data from 15 commercial dairy farmers who have chosen to share their data for research purposes. This data is available to both researchers at SLU and external researchers.
Although the data cannot be published openly, it must be archived. Anna therefore ensured that the datasets from the project were archived at SLU.
– As researchers at a government agency, we are required to archive our data. For example, if there is a suspicion of fraud, it should be possible to access the data, but in this case a confidentiality assessment is required as we have agreements with Växa and the farms. From SLU’s archive, we have also received an identifier, a so-called handle, which can be used to reference metadata, for example in an article, along with information that access to the data in this case may be restricted.
In the first study, Anna did much of the data processing in Excel and statistical analysis in R, but after taking statistics courses, she began to use R more.
– At one of the first sessions, they asked who would be using SAS and one person put their hand up. Then they asked who would be using R and everyone else put their hand up, so I thought that must be the future and I’d better learn it too. That was my first encounter with R and it was brilliant. I’m very grateful for that.
In the second study, Anna also used R for data processing, for example to aggregate data when she was preparing her dataset. She also included R scripts in the metadata publications.
– It feels very open and reproducible when you have an R script that does everything. It becomes clear that someone else can do the same. It took a lot of time because I was new to R, but now I’m getting a lot out of it.
When she was learning R, Anna used ChatGPT for help.
– If you get an error message in R, instead of spending ages googling, you can ask ChatGPT and get a fully corrected piece of code. But as is always the case with an AI chatbot, you can never fully trust the result. You need the expertise to tell whether it’s missed something, or things turned out as you intended.
Anna also points out that there are other drawbacks to AI services.
– For example, it uses a lot of energy. And you have to be careful. You can’t, for instance, enter sensitive data because there’s a risk it will be collected and end up in the wrong hands. But when it comes to pure code, I have no problem with it, because that’s the methodology. I think it’s only good if it’s shared and might be of use to someone else.
Anna is now working on a new project where she is looking at slightly different aspects of the same data, and she feels that the time she spent on data management has been beneficial.
– I have the data structure in place and a data description I can refer to, which explains how the data was collected. I also find the documentation useful. You tend to think that of course you’ll remember, but then two weeks later, or even just the next day, you still wonder how you did something.
Interview: Hanna Östholm. Text: Simon Hallstan