arXiv is now an independent nonprofit! Learn more
License: arXiv.org perpetual non-exclusive license
arXiv:2305.03546v2 [eess.IV] 22 Sep 2023

Breast Cancer Immunohistochemical Image Generation:
a Benchmark Dataset and Challenge Review

Chuang Zh    Shengjie Liu    Zekuan Yu    Feng Xu    Arpit Aggarwal    Germán Corredor    Anant Madabhushi    Qixun Qu    Hongwei Fan    Fangda Li    Yueheng Li    Xianchao Guan    Yongbing Zhang    Vivek Kumar Singh    Farhan Akram    Md. Mostafa Kamal Sarker    Zhongyue Shi    Mulan Jin Thanks: The first two authors contributed equally to this work. Thanks: This work was supported in part by National Key R&D Program of China (2021ZD0109800), and in part by the National Natural Science Foundation of China (81972248). Thanks: C. Zhu and S.J. Liu are with the School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing 100876, China. Thanks: Z.K. Yu is with Academy for Engineering and Technology, Fudan University, Shanghai, China. Thanks: F Xu, Z.Y. Shi and M.L. Jin are with Beijing Chaoyang Hospital, Capital Medical University Beijing, China. Thanks: Corresponding authors: Zekuan Yu and Feng Xu (yzk@fudan.edu.cn; drxufeng@mail.ccmu.edu.cn) Thanks: Ethics committee/IRB of Beijing Chao-Yang Hospital, Capital Medical University gave ethical approval for this work.
Abstract

For invasive breast cancer, immunohistochemical (IHC) techniques are often used to detect the expression level of human epidermal growth factor receptor-2 (HER2) in breast tissue to formulate a precise treatment plan. From the perspective of saving manpower, material and time costs, directly generating IHC-stained images from Hematoxylin and Eosin (H&E) stained images is a valuable research direction. Therefore, we held the breast cancer immunohistochemical image generation challenge, aiming to explore novel ideas of deep learning technology in pathological image generation and promote research in this field. The challenge provided registered H&E and IHC-stained image pairs, and participants were required to use these images to train a model that can directly generate IHC-stained images from corresponding H&E-stained images. We selected and reviewed the five highest-ranking methods based on their PSNR and SSIM metrics, while also providing overviews of the corresponding pipelines and implementations. In this paper, we further analyze the current limitations in the field of breast cancer immunohistochemical image generation and forecast the future development of this field. We hope that the released dataset and the challenge will inspire more scholars to jointly study higher-quality IHC-stained image generation.

Index Terms: 
Breast cancer, pathology image dataset, immunohistochemical image generation, image-to-image translation, grand challenge.

I Introduction

According to data [1] released by the International Agency for Research on Cancer (IARC), female breast cancer has surpassed lung cancer as the most commonly diagnosed cancer in 2020, with an estimated 2.3 million new cases. Early determination of the type and stage of breast cancer is crucial to the formulation of treatment plans and the prognosis of patients.

The current diagnosis of breast cancer is based on the pathological tissue stained with Hematoxylin and Eosin (H&E) as the gold standard. Surgeons take a piece of tissue from the lesion area of the patient, which undergoes a series of procedures including slice preparation and staining, to finally make a pathological slide available for observation. The pathologist then observes the slice under a microscope and provides a diagnosis. An H&E-stained slice is shown in Fig. 1(a)).

For patients diagnosed with breast cancer, specific protein testing is often required to further evaluate the tumor. For example, the state (positive or negative) of human epidermal growth factor receptor-2 (HER2), needs to be identified for breast cancer, as the HER2 state is a helpful marker for therapy decision making [2].

If a patient tests positive for HER2, doctors will administer targeted drug therapy. Timely targeted therapy can increase the survival chance of HER2-positive patients to a level similar to those of HER2-negative patients. Experts recommend that all patients diagnosed with invasive breast cancer undergo HER2 testing to significantly improve the treatment recommendations and decisions [3].

Refer to caption

(a) An example of H&E slice.

Refer to caption

(b) An example of IHC-stained slice.

Fig. 1: Visualization of an H&E-stained slice and the corresponding immunohistochemical (IHC) stained slice.
Refer to caption

(a) IHC 0

Refer to caption

(b) IHC 1+

Refer to caption

(c) IHC 2+

Refer to caption

(d) IHC 3+

Fig. 2: Visualization of different HER2 expression levels. Generally, the cell membrane is stained darker with increasing HER2 expression levels.

The normal method for evaluating HER2 expression levels is to interpret the pathological images stained by immunohistochemical (IHC) technique. Specifically, an additional tissue section is taken from the patient’s pathological tissue for IHC staining (an IHC-stained slice is shown in Fig. 1(b)), and pathologists determine the level of HER2 expression based on the staining pattern of the cell membrane in the section. According to clinical oncology practical guideline of American [4], the interpretation rules for immunohistochemical images of breast cancer are as follows: IHC 0, no staining is observed or membrane staining that is incomplete and is faint/barely perceptible and in 10%\leq 10\% of tumor cells (Fig. 2(a)); IHC 1+, incomplete membrane staining that is faint/barely perceptible and in >10% of tumor cells (Fig. 2(b)); IHC 2+, weak to moderate complete membrane staining observed in >10% of tumor cells (Fig. 2(c)); IHC 3+, circumferential membrane staining that is complete, intense, and in >10% of tumor cells (Fig. 2(d)).

The evaluation of HER2 expression is critical to the formulation of follow-up treatment plans for breast cancer. However, it is expensive to conduct HER2 evaluation through the preparation of an IHC-stained slice. Moreover, a single IHC-stained section may not be able to comprehensively assess the level of HER2 expression in tumor tissue. For instance, in Fig. 3, if the HER2 expression level is diagnosed as 2+, obtaining a new tissue sample for IHC staining may be necessary. This additional preparation requirement for IHC-stained slices further increases the costs associated with labor and materials.

The question that arises is whether it is possible to synthesize an IHC-stained image from an H&E-stained image, and thus avoid the expensive IHC staining. With the advancement of deep learning, many intelligent applications in the medical field have emerged, such as tumor cell classification [5, 6, 7], tumor segmentation [8, 9, 10], and pathology image staining normalization [11, 12, 13], etc. The powerful capabilities demonstrated by deep learning in the above-mentioned fields make us hope that it can also have the ability to directly generate IHC-stained pathological images. Successful exploration of this technology would help save considerable human and financial costs associated with IHC-stained slice preparation. At the same time, to mitigate potential inaccurate HER2 detection stemming from tumor heterogeneity [14], IHC-stained image generation could easily be performed on multiple H&E-stained tumor tissue sections from invasive breast cancer patients. Besides, this study will also help us understand and interpret what information in H&E-stained images is relevant to HER2 assessment in the future.

Fig. 3: Algorithm for evaluation of human epidermal growth factor receptor 2 (HER2) protein expression by immunohistochemistry (IHC) assay of the invasive component of a breast cancer specimen.

The dataset is a crucial factor in studying H&E to IHC-stained image translation, and we have successfully collected and established the Breast Cancer Immunohistochemical (BCI) 11 1 https://bupt-ai-cz.github.io/BCI_for_GrandChallenge, a paired H&E to IHC-stained image translation dataset. In this dataset, H&E-stained images and IHC-stained images are already achieved structure-level alignment.

The dataset provides a foundation for the research of the IHC-stained image generation algorithms. Based on the BCI dataset, we hosted a challenge22 2 https://bci.grand-challenge.org for the generation of breast cancer IHC-stained images, in which participants were required to train an IHC-stained image generation algorithm and submit the generated IHC-stained images. The challenge attracted over 500 registrations and received a total of 75 submissions.

II Related Work

II-A Computer-Aided Diagnosis of Pathology

Deep learning has been widely used in many computer vision tasks such as image classification [15], semantic segmentation [16], and object detection [17]. The extension of the above technologies in the field of pathology has also become a hot research topic. The current applications of deep learning in pathological image analysis include tumor detection and classification, tumor segmentation, cell detection and counting, etc.

The pioneering work [18] gave a series of benchmarks for pathological image-based detection, segmentation and recognition based on Convolutional Neural Networks (CNN). These benchmarks include nuclei segmentation, epithelium segmentation, tubule segmentation, Invasive Ductal Carcinoma (IDC) segmentation, lymphocyte detection, mitosis detection, and lymphoma sub‑type classification. Inspired by the above work and driven by demand, a large number of researches on classification [19, 20] and segmentation [21, 22, 23] of pathological images based on deep learning have emerged. Work [19] designed a transformer-based Multiple Instance Learning (MIL) framework, which can effectively deal with unbalanced/balanced and binary/multiple WSI classification. Work [20] introduced a novel Attention High-order deep Network (AHoNet) by simultaneously embedding attention mechanism and high-order statistical representation into a residual convolutional network and this network can capture more discriminative deep features for breast cancer pathological images. Regarding the detection and segmentation of specific pathological tissue regions, works [21] and [22] used semantic segmentation models (i.e. DeepLabv3 [24], DeepLabv3+ [25]) to generate candidate cancer regions in pathological images for reference by pathologists. Work [23] introduced TissueNet, an extensively annotated tissue image dataset designed for training cell segmentation models. This dataset was employed to train Mesmer, a segmentation model that outperforms previous algorithms in terms of accuracy. Furthermore, some researchers used multi-task learning to analyze pathological images [26] to realize the segmentation and classification tasks in one model. These studies have greatly promoted the development of computer-aided diagnosis. In clinical practice, some mature algorithms have been deployed to the front line, and these algorithms are able to automate repetitive and time-consuming tasks, playing a huge role in reducing the clinical workload of pathologists.

II-B Classification of HER2 Expression Levels

To explore alternatives to HER2 expression level classification without relying on IHC-stained images, some scholars have begun to employ deep learning-based methods to predict HER2 expression levels based on images from other modalities, thereby eliminating the IHC staining step. Xu et al. [27] proposed a DenseNet-based deep learning model using ultrasound images as input to predict HER2 expression and the performance significantly exceeded the traditional texture analysis based on the radiomics model. La Barbera et al. [28] proposed a pipeline that mimics clinician diagnosis: they first employed a cascade of deep neural network classifier for breast cancer screening and then detected the presence of HER2 via MIL. In addition, there are some other studies [29, 30] predict the expression level of HER2 based on H&E-stained images. Work [29] proposed a multi-stage image classification pipeline to realize the classification of breast cancer tumors based on H&E-stained images. Work [30] utilized an Inceptionv3 [31] architecture to predict HER2 status in breast cancer and trastuzumab treatment in HER2-positive samples.

The above studies that predict HER2 expression levels based on H&E-stained images demonstrate that such images contain some information regarding HER2 expression levels. However, the output of these classification algorithms is HER2 positive probability, and the classification model is a black box, which limits the interpretability of the model’s HER2 expression level predictions.

II-C Image-to-Image Translation

Image translation algorithms can be divided into supervised image translation algorithms and unsupervised image translation algorithms according to the form of supervision. For supervised image translation models, pixel-level aligned image pairs are required for supervision during the training phase. Some pioneering works [32, 33] have implemented supervised translation of natural images. Image translation algorithms have significant application value in the field of pathology. Recently, work [34] proposed the PyramidPix2pix model to constrain the generated images on multiple scales and achieved state-of-the-art (SOTA) on the pathology image translation dataset. The impressive work [35] employed the DeepLIIF framework, which is capable of converting IHC images into more informative and higher-cost Multiplex Immunofluorescence (mpIF) images. Different from the above supervised image translation algorithms, unsupervised image translation models [36, 37, 38, 39] are trained with unaligned image pairs and this type of algorithms realize changing the image style mainly by adversarial learning. In the field of medical image analysis, unsupervised image translation algorithms are primarily employed for the staining normalization of pathological images [40, 11, 41] and data augmentation [42, 43, 44, 45, 46, 47].

In this study, we aim to leverage image-to-image translation technology to generate IHC-stained images from corresponding H&E-stained images. This approach would enable us to better visualize the expression of HER2 in tumor tissues, rather than directly predicting the HER2 status.

III Dataset Construction

As a key factor to improve the performance of deep learning models, many open-source datasets have been widely applied in computer vision tasks, such as ImageNet [48] and MNIST [49] for image classification; COCO [50], Cityscapes [51], and KITTI [52] for semantic segmentation and object detection.

However, there are currently no publicly available datasets for conducting research on generating breast cancer IHC-stained images. Study [53] highlights that a major challenge in the current field of Computational Pathology (CPATH) is the lack of publicly available datasets that truly represent clinical practice. Therefore, we introduced the BCI dataset. This dataset serves as a valuable resource for investigating the conversion of H&E-stained breast cancer tissue images into IHC-stained images. The forthcoming section will comprehensively outline the construction methodology employed for the development of the BCI dataset.

Refer to caption
Fig. 4: The process of constructing the BCI dataset. First, we prepared H&E-stained and IHC-stained pathological sections from the tumor tissue extracted from the patient’s breast. Then we used the IHC-stained WSI as the fixed image to register the H&E-stained WSI: projection transformation and elastix registration enabled the H&E-stained WSI to be globally and locally aligned with the IHC-stained WSI. Finally, we refined the transformed WSI and cut the H&E-IHC WSI pair to obtain paired H&E-IHC image patches.

III-A Data Construction Process

As denoted in Fig. 4, the data construction process mainly includes the following steps: slice preparation, scanning, projection transformation, elastix registration, image refinement and patch selection.

During the preparation of a pathological slice pair, two layers of sections needed to be continuously cut out from the same tumor tissue for H&E staining and IHC staining, respectively. Therefore, the section for H&E staining was similar in shape to the corresponding section for IHC staining. The prepared pathological slides were then scanned into WSIs. Specifically, we used Hamamatsu NanoZommer S60 (capable of 20×\times magnification) to scan H&E-stained WSI and the corresponding IHC-stained WSI. Due to computing power and memory limitations, we downsampled the original images by reducing their dimensions to half in both width and height in the subsequent processing steps.

Then the downsampled H&E-IHC WSI pairs were aligned through image registration with two operations: projective transformation and elastix registration. The two transformation steps were employed to achieve alignment between the H&E-stained WSI and corresponding IHC-stained WSI in terms of global contour and internal details, respectively.

III-B Image Registration

To achieve the alignment of H&E-stained and IHC-stained images, we implemented the following two-step registration process:

Projective Transformation. First, we took the IHC-stained image as a reference and performed a projective transformation on the corresponding H&E-stained image. Projective transformation requires no less than 4 selected one-to-one correspondences in the image to be transformed (H&E) and the reference image (IHC). Projective transformation can then perform operations such as translation, scaling, rotation, beveling, and perspective distortion on the image to be transformed so that the H&E-stained image and the IHC-stained image can be initially aligned on the outline.

Elastix Registration. After completing the projective transformation of the H&E-stained image, there were still some misalignments between the H&E and IHC-stained images. We therefore iteratively registered H&E and IHC-stained images using the medical image registration software elastix to further remove these misalignments. In the specific implementation process, since the computer memory cannot carry the high-resolution WSI registration calculation, we divided the H&E and IHC-stained images into 16 parts for registration respectively, and then the registered parts were spliced back into WSI according to the original positional relationship.

The visualization results of the projection transformation and elastix registration are compared in Fig. 6.

III-C Post-Processing

During the registration process, some operations in projective transformation (e.g. the rotation and scaling of the image) left some black areas at the edges of the WSI. At the same time, in the process of elastix registration, in order to align the internal content of the 16 image parts, the edges also moved to their inside resulting in black borders. We implemented an image refinement to address the above problems by filling the black area with surrounding pixels. Finally, we segmented the WSIs into square patches with a side length of 1024 pixels and filtered out regions that did not contain tumor tissue or that were not aligned by the two-step registration procedure.

IV Challenge Setup

IV-A Aims and Tasks

This challenge aims to advance research on generating immunohistochemical images of breast cancer. The generated IHC-stained images should contain accurate HER2 expression information for direct interpretation by physicians. The use of a deep learning model for generating immunohistochemical images of breast cancer has the potential to save time, manpower, and material resources by eliminating the need for IHC-stained section preparation. Participants need to use pairs of H&E and IHC-stained images in the training set to train an image translation model. During the prediction stage, the model must generate IHC-stained images using only H&E-stained images as input. Furthermore, we annotated the extracted image patches using the HER2 expression information (0/1+/2+/3+) at the WSI level, which has been interpreted by pathologists. This label information is only available for model training and not for model prediction. Using label information is optional, but if a participant uses label information in their method, they should indicate it in the comments when submitting their results.

IV-B Data Information

Our dataset contains 4872 pairs of aligned H&E-IHC pathology image patches, which come from the WSIs of more than 300 patients. The HER2 expression levels of these patients encompass four grades: 0, 1+, 2+, and 3+. Fig. 5 illustrates the distribution of HER2 expression levels of the image pairs (We utilize the HER2 expression level of the WSI to represent that of the corresponding image pair). The dataset used in this challenge consists of 3396 pairs of the training set images, 500 pairs of the validation set images, and 977 pairs of test set images. The three subsets are obtained by randomly partitioning these 4,872 pairs of images. Among them, the H&E-IHC image pairs in the training set and validation set are all open to the participants, and for the testing set, only the H&E-stained images are open to the participants. The validation set images are only for participants to test and optimize the performance of their models, and can not be used for model training. The test set images are used for challenge evaluation and to get the final ranking.

Refer to caption
Fig. 5: The distribution of HER2 expression levels within the dataset.
Refer to caption
Fig. 6: Visual comparison of projection transformation results (the top row) and elastix registration results (the bottom row). By overlapping the transformed H&E image with the corresponding IHC-stained image, it can be seen that elastix registration can improve the coincidence of the two images.

IV-C Evaluation

We used Peak Signal to Noise Ratio (PSNR) and Structural Similarity (SSIM) as the metrics to evaluate the quality of the generated images. PSNR is based on the error between the corresponding pixels of two images and is the most widely used objective evaluation index.

PSNR(x,y)=10×log10(281)2MSE,PSNR(x,y)=10{\times}\log_{10}{\frac{{(2^{8}-1)}^{2}}{MSE}}, (1)
MSE(x,y)=1mni=0m1j=0n1[x(i,j)y(i,j)]2,MSE(x,y)=\frac{1}{mn}{\textstyle\sum_{i=0}^{m-1}\sum_{j=0}^{n-1}\left[x\left(i,j\right)-y(i,j)\right]^{2}}, (2)

where xx and yy denote the generated IHC-stained image and the corresponding ground truth, mm and nn denote the width and height of the image.

However, the evaluation result of PSNR may be different from the evaluation result of the Human Visual System (HVS). Therefore, we also used SSIM, which comprehensively measures the differences in image brightness, contrast, and structure.

SSIM(x,y)=(2μxμy+C1)(2σxσy+C2)(μx2+μy2+C1)(σx2σy2+C2),SSIM(x,y)=\frac{(2\mu_{x}\mu_{y}+C_{1})(2\sigma_{x}\sigma_{y}+C_{2})}{(\mu_{x}^{2}+\mu_{y}^{2}+C_{1})(\sigma_{x}^{2}\sigma_{y}^{2}+C_{2})}, (3)

where μx\mu_{x} and σx\sigma_{x} denote the mean and standard deviation of the generated IHC-stained image, μy\mu_{y} and σy\sigma_{y} denote the mean and standard deviation of the ground truth, C1C_{1} and C2C_{2} are constants.

The final ranking of the challenge was calculated by the weighted average of the participants’ SSIM ranking and PSNR ranking:

RFinal=0.4×RPSNR+0.6×RSSIM,R_{Final}=0.4\times R_{PSNR}+0.6\times R_{SSIM}, (4)

where RFinalR_{Final} determines the final ranking (smaller RFinalR_{Final} means higher ranking), RPSNRR_{PSNR} denotes the rank of PSNR and RSSIMR_{SSIM} denotes the rank of SSIM.

V Methods

We have summarized the method descriptions submitted by five top-ranked teams in Table I. Among these teams, three employed fully supervised image translation models (arpitdec5, Just4Fun, vivek23), while the other two utilized weakly supervised image translation models (lifangda02, stan9). With regard to the utilization of supplementary information, teams Just4Fun and stan9 incorporated WSI-level category labels (0, 1+, 2+, 3+) during the training of their respective models.

TABLE I: Brief Comparison of Top-Ranked Participating Methods.
Team Basic architecture HER2 expression level used? Supervision
arpitdec5 Pyramid Pix2pix Full supervision
Just4Fun Self-developed Full supervision
lifangda02 CUT Weak supervision
stan9 WeCREST Weak supervision
vivek23 Pix2pix Full supervision

V-A arptidec5:

The solution of team arptidec5 was built based on the framework of Pyramid Pix2pix in the BCI [34]. In contrast to work [34], they filtered the image pairs in the dataset and performed downsampling during the data preprocessing stage. Specifically, before the images were used as input to the model for training, they went through a quality control process to ensure images with artifacts, cracked tissue, or blurriness were removed from the training process [54]. Around 10% of the images among the train set were not used for training. Then the images were resized from 1024×\times1024 to 256×\times256 and were used as input to the model. The output images obtained from the model were subsequently resized from 256×\times256 to 1024×\times1024 to match the original input dimensions. The training and inference process of team arptidec5 is illustrated in Fig. 7.

Refer to caption
Fig. 7: During the training stage, the team arptidec5 used an open-source pathology image control software to filter out some images that were not suitable for training. Then they used downsampled H&E-IHC image pairs to train a Pyramid Pix2pix model. During the inference stage, they made predictions based on the downsampled H&E-stained images and then upsampled the generated IHC-stained images to the original resolution to obtain the final results.

V-B Just4Fun:

Team Just4Fun proposed a level-aware BCIStainer to learn HER2 expression levels from H&E-stained slices IheI_{he}, and meanwhile, translate IheI_{he} to IHC-stained slices I^ihc\hat{I}_{ihc}. The goals of BCIStainer are: (1) keeping consistency between HER2 expression levels of I^ihc\hat{I}_{ihc} and the ground truth IihcI_{ihc}; (2) generating similar content of I^ihc\hat{I}_{ihc} as IihcI_{ihc}; (3) making I^ihc\hat{I}_{ihc} and IheI_{he} be structurally similar.

The architecture of HER2 expression level-aware BCIStainer (G)(G) is shown in Fig. 8. GencG_{enc} encodes the input IheI_{he} from high resolution (1024×\times1024) to a feature map of size 128×\times128. GclsG_{cls} outputs latent features S^he\hat{S}_{he} from encoded IheI_{he}, and then S^he\hat{S}_{he} is used for predicting HER2 expression y^\hat{y} by a linear classification header. GstainerG_{stainer} shown in Fig. 8 applies S^he\hat{S}_{he} as the condition in the weight-demodulated layer [55] to perform the image-to-image translation process. The participants also added the parameters-free attention layer SimAM [56] in GstainerG_{stainer} to enhance the salient features of cell structure or HER2 expressions, such as cell edges, cell nucleus, and dark regions of high-level HER2 expression. GstainerG_{stainer} consists of 9 basic blocks. At the end of each basic block, they added a residual skip connection before the output. A convolutional layer is followed by GstainerG_{stainer} and outputs the low-resolution prediction I^ihclow\hat{I}_{ihc}^{low}. The final GdecG_{dec} recovers translated feature maps to the original image space and outputs predicted IHC-stained slices I^ihc\hat{I}_{ihc}.

The following is a detailed description of the loss functions used in this framework, including level loss, content loss, and adversarial loss.

Refer to caption
Fig. 8: Team Just4Fun used the architecture of generative adversarial network as a whole, in which the encoder of BCIStainer extracts features from the H&E-stained image for classification, and the classification information, in turn, guides the generation of IHC-stained images. For the generated IHC-stained images, the participants used Mean Absolute Error (MAE), Structural Similarity (SSIM), and Cosine Similarity (CSIM) loss to constrain.

Level Loss. Level loss is the combination of multiple-class focal loss LgfocalL_{gfocal} [57] and cosine similarity loss LcsimL_{csim} in sample level. Because of the imbalanced multiple-class dataset, they used a LgfocalL_{gfocal} to train the classifier in BCIStainer with predicted level y^\hat{y} and the ground truth level yy. Thus S^he\hat{S}_{he} is able to conduct guidance to GstainerG_{stainer}. LgfocalL_{gfocal} can be calculated by the following equations:

Lgfocal=1Nn=1Nαn(1pt,n)γlog(pt,n),L_{gfocal}=\frac{1}{N}\sum_{n=1}^{N}-\alpha_{n}(1-p_{t,n})^{\gamma}log(p_{t,n}), (5)
pt,n={pn,if y=n,1pn,otherwise,p_{t,n}=\begin{cases}p_{n},&\text{if $y=n$},\\ 1-p_{n},&\text{otherwise},\end{cases} (6)
pn=ey^ni=1Ney^i,p_{n}=\frac{e^{\hat{y}_{n}}}{\sum_{i=1}^{N}e^{\hat{y}_{i}}}, (7)

where NN is the number of expression levels, nn is the nnth level, ii is the iith level, pn[0,1]p_{n}\in[0,1] is the predicted probability for the class with label nn, α\alpha is a vector of weight coefficients of all levels, (1pt,n)γ(1-p_{t,n})^{\gamma} is the modulating factor to reshape the loss function to down-weight easy examples, and γ\gamma is a tunable parameter to smoothly adjusts the modulating factor [57]. To compute LcsimL_{csim}, they firstly trained a classifier as comparator CC, only using the real IHC-stained slices IihcI_{ihc} and expression levels yy with a multiple-class focal loss LcfocalL_{cfocal} in the same formulation as (5) to (7). CC generates representations S^ihc\hat{S}_{ihc} and SihcS_{ihc} by inputting I^ihc\hat{I}_{ihc} and IihcI_{ihc}. LcsimL_{csim} measures the similarity between I^ihc\hat{I}_{ihc} and IihcI_{ihc} in sample level as (8). Lower LcsimL_{csim} reveals higher consistency in expression levels of I^ihc\hat{I}_{ihc} and IihcI_{ihc}.

Lcsim=1S^ihc·SihcS^ihc·Sihc.L_{csim}=1-\frac{\hat{S}_{ihc}\textperiodcentered S_{ihc}}{\left\|\hat{S}_{ihc}\right\|\textperiodcentered\left\|S_{ihc}\right\|}. (8)

Content Loss. Content loss is composed of Mean Absolute Error (MAE) loss LmaeL_{mae} and structure similarity loss LssimL_{ssim}. LmaeL_{mae} is used to measure content consistency between full resolution I^ihc\hat{I}_{ihc} and IihcI_{ihc}, and also in the low resolution I^ihclow\hat{I}_{ihc}^{low} and IihclowI_{ihc}^{low}. LmaeL_{mae} is shown as (9):

Lmae=IihcI^ihc1+IihclowI^ihclow.L_{mae}=\left\|I_{ihc}-\hat{I}_{ihc}\right\|_{1}+\left\|I_{ihc}^{low}-\hat{I}_{ihc}^{low}\right\|. (9)

Structure similarity comprehensively measures the differences between images in brightness, contrast, and structure. They applied structure similarity as a loss function to train the generator directly as (10):

Lssim=1SSIM(I^ihc,Iihc).L_{ssim}=1-SSIM(\hat{I}_{ihc},I_{ihc}). (10)

Adversarial Loss. Adversarial loss LGANL_{GAN} is a multiple-scale version of PatchGan from Pix2pixHD [33]. LGAN1024L_{GAN}^{1024} is computed by I^ihc\hat{I}_{ihc} and IihcI_{ihc} in full resolution 1024×\times1024. LGAN512L_{GAN}^{512} is computed in the same way as LGAN1024L_{GAN}^{1024}, but using resized I^ihc\hat{I}_{ihc} and IihcI_{ihc} in resolution 512×\times512. LGANL_{GAN} is the mean of LGAN1024L_{GAN}^{1024} and LGAN512L_{GAN}^{512} as (11).

LGAN\displaystyle L_{GAN} =argminGmaxD(LGAN1024(I^ihc,Iihc)CLOSE\displaystyle=arg\min_{G}\max_{D}(L_{GAN}^{1024}(\hat{I}_{ihc},I_{ihc}) (11)
OPEN+LGAN512(I^ihc,Iihc))×0.5.\displaystyle+L_{GAN}^{512}(\hat{I}_{ihc},I_{ihc}))\times 0.5.

Overall Loss. Overall loss is constructed as (12):

L\displaystyle L =λgfocalLgfocal+λcsimLcsim+λmaeLmae\displaystyle=\lambda_{gfocal}L_{gfocal}+\lambda_{csim}L_{csim}+\lambda_{mae}L_{mae} (12)
+λssimLssim+λGANLGAN,\displaystyle+\lambda_{ssim}L_{ssim}+\lambda_{GAN}L_{GAN},

where λgfocal\lambda_{gfocal}, λcsim\lambda_{csim}, λmae\lambda_{mae}, λssim\lambda_{ssim} and λGAN\lambda_{GAN} are weighting-parameters for corresponding losses.

Team Just4Fun has open-sourced their code at github: https://github.com/quqixun/BCIStainer.

V-C lifangda02

Team lifangda02 summarized the challenging aspects of H&E-to-IHC translation into two points:

(1) Inconsistencies in the H&E-IHC pairs. Since re-staining a slice is physically infeasible, a matching pair of H&E-IHC slices are taken from two depth-wise consecutive cuts of the same tissue and scanned separately. This inevitably prevents pixel-perfect image correspondences due to morphology inconsistency and alignment error. The former is inherent to the fact that the image pair is from separate cuts and their preparation routines might differ. The latter is only exacerbated by the former in the image registration process.

(2) Reproducing the diagnosis-critical characteristics. In this challenge, IHC staining highlights tissue regions with a positive HER2 expression with a brownish color. The higher the HER2 expression is, the darker the brown and the higher the contrast against the benign tissue regions. Therefore, correctly reflecting the HER2 expression levels in the generated IHC-stained images is a huge challenge, especially given the much lower contrast levels between the malignant and benign regions in the H&E-stained images. Additionally, doing so accurately and in a visually discriminative manner is of the core interest of H&E-to-IHC translation.

To address the first challenge, the participants approached the problem of H&E-to-IHC stain transfer from the perspective of “weakly” supervised image-to-image translation. The solution was built on top of the Contrastive Unpaired Translation (CUT) framework by [58]. They augmented the CUT framework with a novel paired contrastive loss, aimed to mitigate the inconsistencies in the H&E-IHC image pairs. To further partially address the second challenge, they designated the discriminator to classify the HER2 level as an auxiliary task.

The CUT framework ensures the content is consistent in the generated image by maximizing the mutual information between input and output. This is implemented by minimizing a patch-based InfoNCE contrastive loss, which aims to learn an embedding that associates corresponding patches to each other, while disassociating them from others. Given a query (a patch) in the output image, the positive is the corresponding patch and the negatives are noncorresponding patches, both from the input image. This loss is denoted as LNCEL_{NCE}. For more details, please refer to (3) in work [58].

The innovative contribution of team lifangda02 is the introduction of paired InfoNCE contrastive loss LpNCEL_{pNCE}, which extends LNCEL_{NCE} to paired images especially to combat the inconsistencies in H&E-IHC image pairs. More specifically, given an output patch as query, the corresponding IHC-stained patch is designated as the positive and the noncorresponding patches are designated as the negatives. Then they used the same InfoNCE-based formulation for LpNCEL_{pNCE}.

The key intuition behind LpNCEL_{pNCE} is that it can be seen as a soft image reconstruction learning criteria. Instead of using a predefined loss term that may not work well on inconsistent ground truth pairs, LpNCEL_{pNCE} punishes dissimilarities between the query and the positive in a learned latent space. Therefore, owing to this adaptiveness, LpNCEL_{pNCE} is more robust towards noisy supervision.

The participants used the resnet-9-blocks generator architecture with the loss function in (13):

G\displaystyle G^{*} =argminGmaxDLGAN+10×LNCE+10×LpNCE\displaystyle=arg\min_{G}\max_{D}L_{GAN}+10{\times}L_{NCE}+10{\times}L_{pNCE} (13)
+2×Ldiscls+20×Lmultiscale,\displaystyle+2{\times}L_{dis-cls}+20{\times}L_{multi-scale},

where LdisclsL_{dis-cls} is the auxiliary cross-entropy-based classification loss using the discriminator and LmultiscaleL_{multi-scale} is the multi-scale image reconstruction loss as introduced in work [34].

An extended version of the method by team lifangda02 can be found in [59], where the authors further extended the LpNCEL_{pNCE} loss to adaptively learn from H&E-IHC image pairs that are more consistent.

Team lifangda02 has open-sourced their code at github: https://github.com/lifangda01/AdaptiveSupervisedPatchNCE.

V-D stan9:

Supervised methods are generally the best methods if paired datasets are available in image translation. However, it is difficult to obtain well paired H&E-IHC images. Unsupervised methods can carry out image translation using unpaired datasets, while the results are less than satisfactory. Therefore, weakly supervised learning may be a better way for H&E to IHC-stained image translation. The method submitted by team stan9 is based on WeCREST [60], which is a weakly supervised deep generative network for style transformation. The backbone of their method is consistent with U-GAT-IT [61]. In each iteration of the training, the input images were sampled according to the sampling rule, which is determined by (14):

Qi=1+cov(Si,Ti)σSiσTij=1NHi·Hj,Qi=\frac{1+\frac{cov(S_{i},T_{i})}{\sigma_{S_{i}}\sigma_{T_{i}}}}{\sqrt{{\textstyle\sum_{j=1}^{N}}H_{i}\textperiodcentered H_{j}}}, (14)

where σSi\sigma_{S_{i}} and σSi\sigma_{S_{i}} are the standard deviations of the source image SiS_{i} and target image TiT_{i}, cov()cov(\cdot) is the covariance, HiH_{i} is a vector of the normalized image histogram for image ii, and NN is the total number of images. The discriminator can classify images into N+1N+1 classes, including NN classes in the pool of real image styles and a fake image style. In addition, the existing methods only consider style transformation but ignore the positive/negative consistency. It is hard for the network to identify cancer areas and the colors of generated images are incorrect sometimes. Therefore, the participants added an auxiliary classifier to the discriminator. The role of the classifier is to classify images according to the HER2 expression status. H&E-stained images are labeled according to the expression status of the corresponding IHC-stained images. As the training goes on, the classification module can make the generator learn pathological characteristics and keep the expression status of the input image and the generated image consistent. The loss function (15) of the generator is composed of adversarial loss LadvL_{adv}, class loss LclassL_{class}, cam loss LcamL_{cam} [61], and cycle loss LcycleL_{cycle}.

LG=λG1Ladv+λG2Lclass+λG3Lcam+λG4Lcycle,L_{G}=\lambda_{G1}L_{adv}+\lambda_{G2}L_{class}+\lambda_{G3}L_{cam}+\lambda_{G4}L_{cycle}, (15)

where λG1\lambda_{G1}, λG2\lambda_{G2}, λG3\lambda_{G3}, λG4\lambda_{G4} are weighting-parameters; adversarial loss LadvL_{adv} has two formats: LadvstL_{adv}^{s\to t} and LadvtsL_{adv}^{t\to s} which are determined by (16) and  17, respectively.

Ladvst=iYiln(Dt(Ti))+Y0ln(Dt(Gst(Si))),L_{adv}^{s\to t}=-\sum_{i}Y_{i}\ln{(D_{t}(T_{i}))}+Y_{0}\ln{(D_{t}(G_{s\to t}(S_{i})))}, (16)

where DtD_{t} is the discriminator of HER2 images, YiY_{i} and Y0Y_{0} are one-hot style labels for image ii and the fake images.

Ladvts=iYiln(Ds(Si))+Y0ln(Ds(Gts(Ti))),L_{adv}^{t\to s}=-\sum_{i}Y_{i}\ln{(D_{s}(S_{i}))}+Y_{0}\ln{(D_{s}(G_{t\to s}(T_{i})))}, (17)

where DsD_{s} is the discriminator of H&E-stained images, YiY_{i} and Y0Y_{0} are one-hot style labels for image ii and the fake images. The loss function of the discriminator ((18)) is composed of adversarial loss LadvL_{adv}, class loss LclassL_{class}, and cam loss LcamL_{cam}.

LD=λD1Ladv+λD2Lclass+λD3Lcam,L_{D}=\lambda_{D1}L_{adv}+\lambda_{D2}L_{class}+\lambda_{D3}L_{cam}, (18)

where λD1\lambda_{D1}, λD2\lambda_{D2}, λD3\lambda_{D3} are weighting-parameters.

Refer to caption
Fig. 9: Team stan9 used a dual image translation model. The discriminators of the two branches classify HER2 expression levels while not only distinguishing whether the images are real or fake. The classification module can make the generator learn pathological characteristics and keep the expression status of the input image and the generated image consistent.

V-E vivek23:

Fig. 10 presents a general overview of the proposed wavelet-based pix2pix model that generates an IHC-stained image from the H&E-stained source image. Team vived23 employed conditional generative adversarial network (cGAN) [32] based on the paired images. The model comprises two sub-networks: a generator that generates a fake or synthetic image, and a discriminator that classifies the generated (fake) image against the corresponding ground truth. The generator network consists of an encoder and a decoder. In the encoder, they incorporated 9 intermediate residual blocks from the ImageNet pre-trained ResNet18 network [15]. Unlike traditional color image-based encoder that works directly on three color channels, they applied the discrete wavelet transform (DWT) to extract four-channel spatial and frequency domain features from the given H&E images [62]. DWT splits the lower and higher-frequency details into sub-bands that help in precisely measuring sharp changes in the input image. In the generator network’s architecture, the first convolutional layer employs 64 filters with a kernel size of 7×\times7 and stride 1, followed by the InstanceNorm and the ReLU activation function [36]. While the encoded features are decoded by the two ConvTranspose2D deconvolutional layers.

Refer to caption
Fig. 10: Team vivek23 used Discrete Wavelet Transform (DWT) to convert RGB images into 4-channel spatial and frequency domain information, then input them into the network. DWT segments low-frequency and high-frequency details in an image, helping to accurately measure sharp changes in the input image.

This team defined ii as the input H&E image, gtgt as the corresponding IHC-stained image, and zz as a random variable. Generator GG and Discriminator DD, generate the output of G(i,z)G(i,z) and D(i,G(i,z))D(i,G(i,z)), respectively. Therefore, the loss function of the generator network GG consists of the Binary CrossEntropy (BCE) loss and L1L1 loss that can be formulated as (19):

LGen(G,D)\displaystyle L_{Gen}(G,D) =Ei,gt,z(logD(i,G(i,z)))\displaystyle=E_{i,gt,z}(-\log{D(i,G(i,z))}) (19)
+λEi,gt,z(L1(gt,G(i,z))),\displaystyle+\lambda E_{i,gt,z}(L_{1}(gt,G(i,z))),

where λ\lambda is an empirical weighting factor that is set to 100. The L1L1 loss helps in reducing the number of false positives and fosters the training process by generating sharp images. The discriminator network utilizes five convolutional layers with a kernel of 4×\times4, and stride of 2×\times2. The early four layers are followed by batch normalization and non-linear leaky ReLU (slope 0.2) activation function. The last layer uses a sigmoid activation function to discriminate between real and fake generated outcomes. The discriminator loss can be defined as (20):

LDis(G,D)\displaystyle L_{Dis}(G,D) =Ei,gt,z(log(D(i,gt)))\displaystyle=E_{i,gt,z}(-\log{(D(i,gt))}) (20)
+λEi,gt,z(log(1D(i,G(i,z)))).\displaystyle+\lambda E_{i,gt,z}(-\log{(1-D(i,G(i,z)))}).

Equation (20) utilizes the BCE loss function for real IHC-stained images against the generated (fake) ones. During training, the discriminator network enforces the generator network in generating a better synthetic IHC-stained image while comparing it with actual ground truth. The discriminator network is not involved in the model’s evaluation phase.

VI Results and Discussion

Refer to caption
Fig. 11: Visualization of invalid submissions. These images are visually blurry. Pathologists cannot give an interpretation of HER2 expression levels based on these images.

We received submissions from a total of 12 teams for our challenge. To evaluate the generated immunohistochemical images submitted by the participants, we invited three doctors to review them. During the review process, we noticed that some of the participants’ submissions contained blurry images where the cellular structure could not be observed, as shown in Fig. 11. Due to their lack of clinical reference value, these submissions were deemed invalid.

Ultimately, we were able to collect 6 valid submissions, and only these submissions were ranked.

VI-A Quantitative Results

Table II shows the PSNR and SSIM metrics of the participants in the challenge. Among them, the team arpitdec5 achieved the highest SSIM and the second highest PSNR, and according to (4), they obtained the highest final ranking. The team Just4Fun obtained the highest PSNR and the second highest SSIM and finally ranked second. Finally, the third to sixth rankings were awarded to teams lifangda02, stan9, guanxianchao, and vivek23, respectively.

TABLE II: Quantitative results of participating teams.
Team Final Rank PSNR(dB)/Rank SSIM/Rank
arpitdec5 1 19.736 / 2 0.574 / 1
Just4Fun 2 22.929 / 1 0.559 / 2
lifangda02 3 17.927 / 5 0.555 / 3
stan9 4 17.959 / 4 0.543 / 4
guanxianchao1 5 19.560 / 3 0.497 / 5
vivek23 6 15.271 / 6 0.493 / 6
  • 1

    Team guanxianchao didn’t submit their method.

Refer to caption

(a) IHC 0

Refer to caption

(b) IHC 1+

Refer to caption

(c) IHC 2+

Refer to caption

(d) IHC 3+

Fig. 12: Visualization results submitted by participating teams. Subfigures (a)-(d) show the generated IHC-stained images at different HER2 expression levels.

VI-B Qualitative Results

Fig. 12 shows the generated IHC-stained images submitted by the participants. In the results submitted by the team Just4Fun, the color depth of the cells is overall consistent with the ground truth, which indicates that their results can accurately reflect the expression level of HER2 to a large extent. The results submitted by the team stan9 can approach the ground truth when HER2 expression levels are 0, 1+, and 2+, but when HER2 is highly expressed (3+), their results cannot accurately reflect the HER2 expression level. The results of the other four teams show similar staining intensity at HER2 expression levels of 0, 1+, 2+, and 3+, and these results cannot accurately reflect the information on HER2 expression levels.

VI-C Discussion

The difficulty of this challenge is how to generate IHC-stained images that can correctly reflect the expression level of HER2 based on H&E images. Most of the participants’ methods have the problem of not being able to identify the high-level expression of HER2 in breast tissue. For example, the coloring degree of the IHC-stained images generated by the teams arpitdec5, lifangda02 and vivek23 almost have the same tone, showing light brown, which is a phenomenon of mode collapse. The images submitted by Just4fun can reflect the expression level of HER2 to a large extent: when the expression level of HER2 is low (0/1+), the generated IHC-stained images are lightly colored, and when the expression level of HER2 is high (2+/3+), the generated IHC-stained images exhibit darker coloration.

The ability of the generated IHC-stained images to accurately reflect the HER2 expression level is highly related to whether the corresponding method uses HER2 expression level information. Team Just4fun made good use of the category information of the original WSIs: their method uses category information to supervise the classification of H&E-stained images (GclassG_{class}), and uses H&E category information to intervene in the generation of IHC-stained images; at the same time, the method also inputs the generated IHC-stained images and the real IHC-stained images into the classification network CC, and uses cosine similarity loss to constrain the category information of the two to be consistent. The team stan9 also used label information. Their methed uses a cycle structure similar to cycleGAN [36], and adds a classification branch to the two discriminators, constraining the two discriminators to agree on the classification of an H&E-stained image and the corresponding IHC-stained image. Therefore, the results submitted by the team stan9 can also generate IHC-stained images with different degrees of coloring, and achieve relatively accurate results in the case of IHC 2+, but still cannot generate correctly stained IHC images in the case of IHC 3+.

VII Conclusion

The results of this breast cancer immunohistochemical image generation challenge indicate that producing high-quality IHC-stained images from H&E-stained images, which accurately depict HER2 expression levels, remains a significant challenge. However, the methods submitted by the participants still provide many novel ideas for IHC-stained image generation.

According to the collected methods, the integration of WSI-level label information can make the generated IHC-stained images avoid mode collapse and correctly reflect the expression level of HER2 to a certain extent. At the same time, the strongly constrained L1 loss used in the fully supervised methods may affect the quality of the generated image, while the weakly supervised or unsupervised methods do not impose pixel-level constraints on the images, so they can better maintain the cell structure in the pathological image. The weakly supervised/unsupervised image translation algorithm has great potential in the field of pathological image generation. In order to promote related research, we will release unaligned H&E-IHC image pairs as an expansion of the BCI dataset. To explore more possibilities of breast cancer pathological image generation, we have opened the post-challenge submission phase, and scholars can still submit their results on Grand Challenge: https://bci.grand-challenge.org.

Acknowledgment

Arpit Aggarwal and Anant Madabhushi are with Biomedical Engineering Department, Georgia Tech and Emory University (e-mail: aagga56@emory.edu, anantm@emory.edu). Anant Madabhushi is also a Research Career Scientist with the Atlanta Veterans Affairs Medical Center.

Germán Corredor is with Biomedical Engineering Department, Emory University (e-mail: gcorred@emory.edu).

Qixun Qu is an independent researcher (e-mail: quqixun@gmail.com).

Hongwei Fan is with Data Science Institute, Imperial College London (e-mail: h.fan21@imperial.ac.uk).

Fangda Li is with the School of Computer and Electrical Engineering, Purdue University (e-mail: li1208@purdue.edu).

Yueheng Li, Xianchao Guan and Yongbing Zhang are with the School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen) (e-mail: 22s051028@stu.hit.edu.cn; 21s051007@stu.hit.edu.cn; ybzhang08@hit.edu.cn).

Vivek Kumar Singh is with the Department of Computer Engineering and Mathematics, Rovira I Virgili University (e-mail: vivekkr.singh90@gmail.com).

Farhan Akram is with the Department of Pathology and Clinical Bioinformatics, Erasmus Medical Center (e-mail: f.akram@erasmusmc.nl).

Md. Mostafa Kamal Sarker is with the Institute of Biomedical Engineering, University of Oxford, Oxford, UK (e-mail: md.sarker@eng.ox.ac.uk).

References

  • [1] H. Sung, J. Ferlay, R. L. Siegel, M. Laversanne, I. Soerjomataram, A. Jemal, and F. Bray, “Global cancer statistics 2020: Globocan estimates of incidence and mortality worldwide for 36 cancers in 185 countries,” CA: a cancer journal for clinicians, vol. 71, no. 3, pp. 209–249, 2021.
  • [2] N. Iqbal and N. Iqbal, “Human epidermal growth factor receptor 2 (her2) in cancers: overexpression and therapeutic implications,” Molecular biology international, vol. 2014, 2014.
  • [3] D. Yamauchi and D. Hayes, “Her2 and predicting response to therapy in breast cancer,” UpToDate, last updated Oct, vol. 4, p. 20122012, 2008.
  • [4] A. C. Wolff, M. E. H. Hammond, K. H. Allison, B. E. Harvey, P. B. Mangu, J. M. Bartlett, M. Bilous, I. O. Ellis, P. Fitzgibbons, W. Hanna, et al., “Human epidermal growth factor receptor 2 testing in breast cancer: American society of clinical oncology/college of american pathologists clinical practice guideline focused update,” Archives of pathology & laboratory medicine, vol. 142, no. 11, pp. 1364–1382, 2018.
  • [5] L. Zhang, L. Lu, I. Nogues, R. M. Summers, S. Liu, and J. Yao, “Deeppap: deep convolutional networks for cervical cell classification,” IEEE journal of biomedical and health informatics, vol. 21, no. 6, pp. 1633–1643, 2017.
  • [6] X. Xie, C.-C. Fu, L. Lv, Q. Ye, Y. Yu, Q. Fang, L. Zhang, L. Hou, and C. Wu, “Deep convolutional neural network-based classification of cancer cells on cytological pleural effusion images,” Modern Pathology, vol. 35, no. 5, pp. 609–614, 2022.
  • [7] G. Pattarone, L. Acion, M. Simian, R. Mertelsmann, M. Follo, and E. Iarussi, “Learning deep features for dead and living breast cancer cell classification without staining,” Scientific reports, vol. 11, no. 1, p. 10304, 2021.
  • [8] M. Havaei, A. Davy, D. Warde-Farley, A. Biard, A. Courville, Y. Bengio, C. Pal, P.-M. Jodoin, and H. Larochelle, “Brain tumor segmentation with deep neural networks,” Medical image analysis, vol. 35, pp. 18–31, 2017.
  • [9] W. Wang, C. Chen, M. Ding, H. Yu, S. Zha, and J. Li, “Transbts: Multimodal brain tumor segmentation using transformer,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part I 24, pp. 109–119, Springer, 2021.
  • [10] Y. Jiang, Y. Zhang, X. Lin, J. Dong, T. Cheng, and J. Liang, “Swinbts: A method for 3d multimodal brain tumor segmentation using swin transformer,” Brain Sciences, vol. 12, no. 6, p. 797, 2022.
  • [11] M. T. Shaban, C. Baur, N. Navab, and S. Albarqouni, “Staingan: Stain style transfer for digital histological images,” in 2019 Ieee 16th international symposium on biomedical imaging (Isbi 2019), pp. 953–956, IEEE, 2019.
  • [12] J.-S. Lee and Y.-X. Ma, “Stain style transfer for histological images using s3cgan,” Sensors, vol. 22, no. 3, p. 1044, 2022.
  • [13] J. C. G. Pérez, D. O. Baguer, and P. Maass, “Staincut: Stain normalization with contrastive learning,” Journal of Imaging, vol. 8, no. 7, 2022.
  • [14] G. Turashvili and E. Brogi, “Tumor heterogeneity in breast cancer,” Frontiers in medicine, vol. 4, p. 227, 2017.
  • [15] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
  • [16] O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Medical Image Computing and Computer-Assisted Intervention–MICCAI 2015: 18th International Conference, Munich, Germany, October 5-9, 2015, Proceedings, Part III 18, pp. 234–241, Springer, 2015.
  • [17] S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,” Advances in neural information processing systems, vol. 28, 2015.
  • [18] A. Janowczyk and A. Madabhushi, “Deep learning for digital pathology image analysis: A comprehensive tutorial with selected use cases,” Journal of pathology informatics, vol. 7, no. 1, p. 29, 2016.
  • [19] Z. Shao, H. Bian, Y. Chen, Y. Wang, J. Zhang, X. Ji, et al., “Transmil: Transformer based correlated multiple instance learning for whole slide image classification,” Advances in neural information processing systems, vol. 34, pp. 2136–2147, 2021.
  • [20] Y. Zou, J. Zhang, S. Huang, and B. Liu, “Breast cancer histopathological image classification using attention high-order deep network,” International Journal of Imaging Systems and Technology, vol. 32, no. 1, pp. 266–279, 2022.
  • [21] Z. Song, S. Zou, W. Zhou, Y. Huang, L. Shao, J. Yuan, X. Gou, W. Jin, Z. Wang, X. Chen, et al., “Clinically applicable histopathological diagnosis system for gastric cancer detection using deep learning,” Nature communications, vol. 11, no. 1, p. 4294, 2020.
  • [22] J. Wang and X. Liu, “Medical image recognition and segmentation of pathological slices of gastric cancer based on deeplab v3+ neural network,” Computer Methods and Programs in Biomedicine, vol. 207, p. 106210, 2021.
  • [23] N. F. Greenwald, G. Miller, E. Moen, A. Kong, A. Kagel, T. Dougherty, C. C. Fullaway, B. J. McIntosh, K. X. Leow, M. S. Schwartz, et al., “Whole-cell segmentation of tissue images with human-level performance using large-scale data annotation and deep learning,” Nature biotechnology, vol. 40, no. 4, pp. 555–565, 2022.
  • [24] L.-C. Chen, G. Papandreou, F. Schroff, and H. Adam, “Rethinking atrous convolution for semantic image segmentation,” arXiv preprint arXiv:1706.05587, 2017.
  • [25] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schroff, and H. Adam, “Encoder-decoder with atrous separable convolution for semantic image segmentation,” in Proceedings of the European conference on computer vision (ECCV), pp. 801–818, 2018.
  • [26] S. Graham, Q. D. Vu, M. Jahanifar, S. E. A. Raza, F. Minhas, D. Snead, and N. Rajpoot, “One model is all you need: multi-task learning enables simultaneous histology image segmentation and classification,” Medical Image Analysis, vol. 83, p. 102685, 2023.
  • [27] Z. Xu, Q. Yang, M. Li, J. Gu, C. Du, Y. Chen, and B. Li, “Predicting her2 status in breast cancer on ultrasound images using deep learning method,” Frontiers in oncology, vol. 12, p. 829041, 2022.
  • [28] D. La Barbera, A. Polónia, K. Roitero, E. Conde-Sousa, and V. Della Mea, “Detection of her2 from haematoxylin-eosin slides through a cascade of deep learning classifiers via multi-instance learning,” Journal of Imaging, vol. 6, no. 9, p. 82, 2020.
  • [29] D. Anand, N. C. Kurian, S. Dhage, N. Kumar, S. Rane, P. H. Gann, and A. Sethi, “Deep learning to estimate human epidermal growth factor receptor 2 status from hematoxylin and eosin-stained breast tissue images,” Journal of pathology informatics, vol. 11, no. 1, p. 19, 2020.
  • [30] S. Farahmand, A. I. Fernandez, F. S. Ahmed, D. L. Rimm, J. H. Chuang, E. Reisenbichler, and K. Zarringhalam, “Deep learning trained on hematoxylin and eosin tumor region of interest predicts her2 status and trastuzumab treatment response in her2+ breast cancer,” Modern Pathology, vol. 35, no. 1, pp. 44–51, 2022.
  • [31] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna, “Rethinking the inception architecture for computer vision,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2818–2826, 2016.
  • [32] P. Isola, J.-Y. Zhu, T. Zhou, and A. A. Efros, “Image-to-image translation with conditional adversarial networks,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1125–1134, 2017.
  • [33] T.-C. Wang, M.-Y. Liu, J.-Y. Zhu, A. Tao, J. Kautz, and B. Catanzaro, “High-resolution image synthesis and semantic manipulation with conditional gans,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 8798–8807, 2018.
  • [34] S. Liu, C. Zhu, F. Xu, X. Jia, Z. Shi, and M. Jin, “Bci: Breast cancer immunohistochemical image generation through pyramid pix2pix,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pp. 1815–1824, June 2022.
  • [35] P. Ghahremani, Y. Li, A. Kaufman, R. Vanguri, N. Greenwald, M. Angelo, T. J. Hollmann, and S. Nadeem, “Deep learning-inferred multiplex immunofluorescence for immunohistochemical image quantification,” Nature machine intelligence, vol. 4, no. 4, pp. 401–412, 2022.
  • [36] J.-Y. Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE international conference on computer vision, pp. 2223–2232, 2017.
  • [37] M.-Y. Liu, T. Breuel, and J. Kautz, “Unsupervised image-to-image translation networks,” Advances in neural information processing systems, vol. 30, 2017.
  • [38] X. Huang, M.-Y. Liu, S. Belongie, and J. Kautz, “Multimodal unsupervised image-to-image translation,” in Proceedings of the European conference on computer vision (ECCV), pp. 172–189, 2018.
  • [39] H.-Y. Lee, H.-Y. Tseng, J.-B. Huang, M. Singh, and M.-H. Yang, “Diverse image-to-image translation via disentangled representations,” in Proceedings of the European conference on computer vision (ECCV), pp. 35–51, 2018.
  • [40] H. Cho, S. Lim, G. Choi, and H. Min, “Neural stain-style transfer learning using gan for histopathological images,” arXiv preprint arXiv:1710.08543, 2017.
  • [41] S. Cai, Y. Xue, Q. Gao, M. Du, G. Chen, H. Zhang, and T. Tong, “Stain style transfer using transitive adversarial networks,” in International Workshop on Machine Learning for Medical Image Reconstruction, pp. 163–172, Springer, 2019.
  • [42] C. Han, Y. Kitamura, A. Kudo, A. Ichinose, L. Rundo, Y. Furukawa, K. Umemoto, Y. Li, and H. Nakayama, “Synthesizing diverse lung nodules wherever massively: 3d multi-conditional gan-based ct image augmentation for object detection,” in 2019 International Conference on 3D Vision (3DV), pp. 729–737, IEEE, 2019.
  • [43] A. Gupta, S. Venkatesh, S. Chopra, and C. Ledig, “Generative image translation for data augmentation of bone lesion pathology,” in International Conference on Medical Imaging with Deep Learning, pp. 225–235, PMLR, 2019.
  • [44] C. Han, L. Rundo, R. Araki, Y. Nagano, Y. Furukawa, G. Mauri, H. Nakayama, and H. Hayashi, “Combining noise-to-image and image-to-image gans: Brain mr image augmentation for tumor detection,” Ieee Access, vol. 7, pp. 156966–156977, 2019.
  • [45] F. Mahmood, R. Chen, D. Borders, G. N. McKay, K. Salimian, A. Baras, and N. J. Durr, “Adversarial u-net with spectral normalization for histopathology image segmentation using synthetic data,” in Medical Imaging 2019: Digital Pathology, vol. 10956, pp. 137–141, SPIE, 2019.
  • [46] S. Liu, Z. Shah, A. Sav, C. Russo, S. Berkovsky, Y. Qian, E. Coiera, and A. Di Ieva, “Isocitrate dehydrogenase (idh) status prediction in histopathology images of gliomas using deep learning,” Scientific reports, vol. 10, no. 1, pp. 1–11, 2020.
  • [47] K. Stacke, G. Eilertsen, J. Unger, and C. Lundström, “Measuring domain shift for deep learning in histopathology,” IEEE journal of biomedical and health informatics, vol. 25, no. 2, pp. 325–336, 2020.
  • [48] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition, pp. 248–255, Ieee, 2009.
  • [49] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.
  • [50] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft coco: Common objects in context,” in European conference on computer vision, pp. 740–755, Springer, 2014.
  • [51] M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benenson, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 3213–3223, 2016.
  • [52] A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? the kitti vision benchmark suite,” in 2012 IEEE conference on computer vision and pattern recognition, pp. 3354–3361, IEEE, 2012.
  • [53] J. Van der Laak, G. Litjens, and F. Ciompi, “Deep learning in histopathology: the path to the clinic,” Nature medicine, vol. 27, no. 5, pp. 775–784, 2021.
  • [54] A. Janowczyk, R. Zuo, H. Gilmore, M. Feldman, and A. Madabhushi, “Histoqc: an open-source quality control tool for digital pathology slides,” JCO clinical cancer informatics, vol. 3, pp. 1–7, 2019.
  • [55] T. Karras, S. Laine, M. Aittala, J. Hellsten, J. Lehtinen, and T. Aila, “Analyzing and improving the image quality of stylegan,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 8110–8119, 2020.
  • [56] L. Yang, R.-Y. Zhang, L. Li, and X. Xie, “Simam: A simple, parameter-free attention module for convolutional neural networks,” in International conference on machine learning, pp. 11863–11874, PMLR, 2021.
  • [57] T.-Y. Lin, P. Goyal, R. Girshick, K. He, and P. Dollár, “Focal loss for dense object detection,” in Proceedings of the IEEE international conference on computer vision, pp. 2980–2988, 2017.
  • [58] T. Park, A. A. Efros, R. Zhang, and J.-Y. Zhu, “Contrastive learning for unpaired image-to-image translation,” in European conference on computer vision, pp. 319–345, Springer, 2020.
  • [59] F. Li, Z. Hu, W. Chen, and A. Kak, “Adaptive supervised patchnce loss for learning h&e-to-ihc stain translation with inconsistent groundtruth image pairs,” arXiv preprint arXiv:2303.06193, 2023.
  • [60] W. Dai, I. H. Wong, and T. T. Wong, “A weakly supervised deep generative model for complex image restoration and style transformation,” 2022.
  • [61] J. Kim, M. Kim, H. Kang, and K. Lee, “U-gat-it: Unsupervised generative attentional networks with adaptive layer-instance normalization for image-to-image translation,” arXiv preprint arXiv:1907.10830, 2019.
  • [62] V. K. Singh, E. Y. Kalafi, S. Wang, A. Benjamin, M. Asideu, V. Kumar, and A. E. Samir, “Prior wavelet knowledge for multi-modal medical image segmentation using a lightweight neural network with attention guided features,” Expert Systems with Applications, vol. 209, p. 118166, 2022.