Aquaculture breeding programs rely on a high number of full-sib familie with large number of offspring, which has facilitated the uptake of genomic selection (GS). These sib-testing schemes yield high prediction accuracies but require genotyping large training populations and many candidates for effective selection. Lowering genotyping costs could broaden GS use for smaller companies while for larger ones, that have already been implementing GS, reducing the individual cost of genotyping would allow the genotyping of additional candidates, resulting in an increased selection intensity and genetic gain. Recent efforts focused on designing low-density (LD) panels combined with genotype imputation to achieve accuracy similar to medium or high-density panels. In this study we compared scenarios to optimize reference populations for accurate LD-genotype imputation in Atlantic salmon. We invetsigated how robust or suboptimal imputation affects prediction accuracy for resistance to cardiomyopathy syndrome (CMS), pancreas disease (PD), and growth (GW).
Materials and Methods
Genotypes, pedigree and phenotypes where obtained from salmons were obtained from MOWI Norway Atlantic salmon breeding programme. All the phenotyped fish and the parents were genotyped with the ThermoFisher Axiom 57K SNP array. After standard quality controls 47,061 SNPs and 1,094 fish for resistance to CMS (gill score), 1,448 for resistance to PD (binary survival) and 2,998 for GW (gutted weight at harvest) were retained. A LD-panel containing 554 pre-established SNPs, was created in silico by masking the genotypes of phenotyped fish (LD-target population). The "missing" genotypes of individuals from the LD-target population were imputed to the high density (HD) using the FImpute software. The impact of the design of the HD-genotyped reference population for imputation was tested in three different scenarios with the HD-genotyped reference population composed of i) all the parents (RPO: Reference Parents Only), ii) the offspring from the validation set only (1/10th of the offspring; RVO: Reference Validation Only) and iii) all parents and the offspring from the validation set (RPV: Reference Parents Validation). The accuracy of imputation was estimated as the Pearson's correlation between the true and the imputed genotype for each individual.
The quality of genomic prediction using pedigree, LD, HD and imputed genotypes was tested in a standard 10-fold cross-valiation approach were (genomic) estimated breeding values ([G]EBV) were computed using the BLUPF90 software with a mixed linear animal BLUP model. Accuracy of prediction was computed as the mean over 10 replicates of the Pearson's correlation coefficient between the (G)EBV and the true phenotypes of fish in the validation group, divided by the square root of the genomic based heritability estimated with the HD-panel.
Results and discussion
Imputation was overall very accurate (>0.8) for all traits. Scenarios that contained HD-genotyped parents and offspring (RPV and RPV) resulted in the highest accuracy while the scenario with only HD-genotyped offspring (RVO) resulted in a significantly lower imputation accuracy. The genotype of idividuals with one or both parents missing from the dataset was badly imputed (correlation <0.5) in scenarios with only parents and the imputation accuracy of those individuals was significantly improved when siblings were added to the HD-reference panel.
For all three traits the accuracy of GEBVs increased significantly with imputation compared to the un-imputed LD-panel. The comparison between the accuracy of prediction obtained from the HD panel and from the imputed panel varied slightly across the traits. For GW, the accuracy of the HD-panel was significantly better than all imputed scenario except for the RPV scenario. For resistance to PD, the predictions with the HD-panel were always significantly more accurate than the predictions obtained with imputed panels. For resistance to CMS, all imputation scenarios were as good as the HD-panel and as good as each other to predict the GEBV of the fish. Those differences may be due to different genetic architecture of the trait with GW being highly polygenic, PD being oligogenic and CMS being monogenic. Interestingly, the difference in the accuracy of imputation obtained between scenarios was not always reflected in the accuracy of genomic prediction. Indeed, the genomic predictions were robust even when genotype imputation was sub-optimal (RVO, Figure 1), althouhg if this scenario resulted in slightly under-dispersed GEBV for GW and CMS.
Conclusion
As aquaculture breeding programmes rely on phenotypes measured on a training population composed of collaterals (full/half-sib) of the selection candidates, an imputation strategy relying on including offspring in an HD-reference panel is particularly suitable for low-cost genomic selection. This strategy could be used to only perform HD-genotyping every two generations, either reducing the cost of genotyping or increasing the selection intensity with minimal decrease in the accuracy of prediction.
Figure 1. Accuracy of imputation and predictions for three traits and different imputation scenarios.
Acknowledgment
This work was funded by the European Union's Horizon 2020 research and innovation programme under the grant agreement No 818367 - AquaIMPACT. The authors were supported by Biotechnology and Biological Sciences Research Council Institute Strategic Grants BBS/E/D/30002275, BBS/E/D/20002172, BBS/E/RL/230001A and BBS/E/RL/230001C to the Roslin Institute.