Introduction
European whitefish (Coregonus lavaretus) is a member of Salmonids used for aquaculture, commercial fisheries, and conservation across northern and central Europe. It is valued for its fast growth, product quality, and adaptability to different production systems (Kankainen et al. 2016). However, the lack of a high-quality reference genome has limited the application of modern genomic tools in selective breeding and management. A major challenge however, is that like other salmonids (Macqueen and Johnston 2014), it has experienced an ancient whole-genome duplication, and parts of the genome still retain duplicated regions (Pokharel et al. 2025) . This makes it difficult to distinguish true biological duplication from technical redundancy during genome assembly. Resolving this issue is important, as incorrect handling of duplicated regions can reduce the usefulness of the genome for downstream applications such as marker discovery and genomic selection. In this study, we aimed to generate a high-quality, chromosome-scale, haplotype-resolved genome assembly for European whitefish and to clarify the extent of duplicated genome content.
Materials and Methods
A mature female European whitefish from the Finnish national breeding programme maintained by Luke (Finland, 2019 year class) was sequenced using PacBio HiFi technology (~152.8 Gb). To improve long-range genome assembly, Omni-C proximity ligation libraries were prepared from fin, liver, and muscle tissues and sequenced on an Illumina NovaSeq (~357 Gb). Genome characteristics, including size, heterozygosity, and ploidy structures, were estimated using k-mer–based approaches including GenomeScope2 (Ranallo-Benavidez et al. 2020), Jellyfish (Mar��ais and Kingsford 2011), and Smudgeplot (Ranallo-Benavidez et al. 2020). Multiple assembly strategies were tested using Hifiasm (Cheng et al. 2021) with different levels of duplicate purging to balance removal of redundant sequences while retaining biologically meaningful duplicated regions. Selected assemblies were scaffolded using Yahs (Zhou et al. 2023) with Omni-C data, and quality was assessed based on contiguity metrics and BUSCO (Manni et al. 2021) gene completeness.
Results
Genome analyses supported a diploid genome structure with substantial retained duplicated content consistent with salmonid whole-genome duplication. Approximately 45–60% of the genome showed signals of incomplete diploid resolution, reflecting retained duplicated regions rather than higher ploidy. Across different assembly strategies, consistent results were obtained for genome size and structure. The final selected assembly achieved chromosome-scale contiguity, with scaffold N50 values of ~68 Mb. The primary assembly size was ~3.09 Gb, with two alternative haplotypes of ~2.89 Gb and ~2.49 Gb. BUSCO completeness excedded 99%, indicating a highly complete assembly.
Discussion
This study provides the first haplotype-resolved, chromosome-scale genome assembly for European whitefish focusing on separating the pure diploid and the duplicated regions. The results support a diploid genome structure with substantial retention of duplicated regions, reflecting the species' evolutionary history rather than technical artifacts. So far, the single-nucleotide polymorphism (SNP) arrays used in modern aquaculture breeding programs are typically limited to the diploid part of the genome. This is a significant issue in a breeding context, as traits encoded in the non-diploid parts of the genome are currently not captured, and entire gene families potentially generated by the duplication event may be unreachable for breeders. The reference genome generated here provides a robust genomic resource for genomic selection, marker development, and population genetic analyses, supporgint both fundamental research and development of genomic tools for European whitefish aquaculture and management.
Acknowledgment
This work has received funding from the Research Council of Finland (contract number: 369785; HiddenGen), Kolarctic Cross-Boarder-Cooperation Programme 2014-2020 (4/2018/095/KO4058; ArctAqua), and the Statutory Services of Natural Resources Institute Finland. CSC-IT Center for Science is acknowledged for computational resources.
References
Cheng, Haoyu, Gregory T. Concepcion, Xiaowen Feng, Haowen Zhang, and Heng Li. 2021. "Haplotype-Resolved de Novo Assembly Using Phased Assembly Graphs with Hifiasm." Nature Methods 18 (2): 170–75.
Kankainen, Markus, Jari Set��l��, Antti Kause, Cheryl Quinton, Susanna Airaksinen, and Juha Koskela. 2016. "Economic Values of Supply Chain Productivity and Quality Traits Calculated for a Farmed European Whitefish Breeding Program." Aquaculture Economics & Management 20 (2): 131–64.
Macqueen, Daniel J., and Ian A. Johnston. 2014. "A Well-Constrained Estimate for the Timing of the Salmonid Whole Genome Duplication Reveals Major Decoupling from Species Diversification." Proceedings of the Royal Society B: Biological Sciences 281 (1778): 20132881.
Manni, Mos��, Matthew R. Berkeley, Mathieu Seppey, and Evgeny M. Zdobnov. 2021. "BUSCO: Assessing Genomic Data Quality and Beyond." Current Protocols 1 (12): e323.
Mar��ais, Guillaume, and Carl Kingsford. 2011. "A Fast, Lock-Free Approach for Efficient Parallel Counting of Occurrences of k-Mers." Bioinformatics 27 (6): 764–70.
Pokharel, Kisun, Daniel Fischer, Terhi Iso-Touru, et al. 2025. "The Genome Assembly of the Farmed European Whitefish Coregonus Lavaretus L. from the Finnish Selective Breeding Programme." BMC Genomic Data 26 (1): 22.
Ranallo-Benavidez, T. Rhyker, Kamil S. Jaron, and Michael C. Schatz. 2020. "GenomeScope 2.0 and Smudgeplot for Reference-Free Profiling of Polyploid Genomes." Nature Communications 11 (1): 1432.
Zhou, Chenxi, Shane A. McCarthy, and Richard Durbin. 2023. "YaHS: Yet Another Hi-C Scaffolding Tool." Bioinformatics 39 (1): btac808.