Dit portaal werkt voor & door studenten. Heb je recent een examen gemaakt? Help je medestudenten en stuur je vragen in!
Examenjaar 2025-2026
Examen PDF
Datum: 2026-08-25
Examen PDF
Datum: 2026-08-25
Examen Vragen
Datum: 2026-08-25
Examen 2025-2026
Bijlagen
📦 Download originele bestanden (ZIP)
#Q1
#Eerst gaje van scratch best een dataframe aanmaken
mydata <- data.frame(
age = rep(c(-1,1), each = 20),
adjuvant = rep(rep(c(1,-1), times = c(5,15)), 2)
)
#rep(c(1,-1), times = c(5,15) dit is voor 1 ggroep dus da gaje nog is moeten repeaten
X <- model.matrix( ~ age + adjuvant + age:adjuvant,
data = mydata)
#b
#dus we weten al de formule van onze variantie covariantie matrix vd coeff
sigma <- 0.2^2*solve(crossprod(X))
#sigma is dus de residual dinges maal die inverse
#hieruit kunnen we dan de SE gaan halen : 2b2 + 2b3
#covariantie van sommige is 0 dus da valt dan al weg
sqrt(2^2*sigma[3,3] + 2^2*sigma[4,4])
#c
m <- matrix(
c(1,1 , 1, 1,
-1, -1, 1, 1,
-1, 1, -1, 1,
1, -1, -1, 1), ncol = 4
)
solution <- c(4.8, 4.8+0.3, 4.3, 4.3+0.6)
solve(m, solution)
#d
dgdata <- cbind(mydata,
Dmut = -0.6)
Xstar <- model.matrix( ~ age + adjuvant + Dmut+ age:adjuvant,
data = dgdata)
round(solve(crossprod(X)) %*% t(X) %*% Xstar, 6)
#lage nummers komen door floating point error dus daarom rounding
#Q2
#data is nie aanwezig dus doe alsof het de juiste resulaten geeft
a <- read.csv("adjuvant.csv",
stringsAsFactors = TRUE)
library(nlme)
mod <- lme(logELU ~ age + treatment + age:treatment,
random = ~ 1|docid/hospital,
data = a)# zie screenshot in word docu want daarin gaje zien hoe je kan weten welke uw random eff zijn
mod <- lme(logELU ~ age + treatment + age:treatment,
random = ~ 1|hospital/docid,
data = a)
summary(mod)
#zo doe je de juiste testen
modred <- lme(logELU ~ age + treatment,
random = ~ 1|hospital/docid,
data = a)
anova(mod,modred)#anova is niet de naam vd test (analysis of variance), de test is een likelihood ratio test
#hieruit zouje dan een waarde krijgen van pvalue maar da nummer is fout want alsje een likelihood ratio
#test doet dan moeje max likelihood ratio doen en niet restricted dus moeje d emethod meegeven
mod <- lme(logELU ~ age + treatment + age:treatment,
random = ~ 1|hospital/docid,
method = "ML",
data = a)
modred <- lme(logELU ~ age + treatment,
random = ~ 1|hospital/docid,
method = "ML",
data = a)
anova(mod,modred)
#Dan krijg je dus de Lwaarde en Pwaarde dieje moe reporten
#eerst dus de plots inladen die meegegeven waren op het examen
residualplot(mod)
ranef(mod)#random effecten eruit krijgen, kijken of ze normaal verdeeld zijn (zowel voor hospitals als dokters in hospitals)
ranef(mod)[[1]] #1ste level eruit krijgen want das hospital en dit is een dataframe wat je dus niet kunt plotten
qqnorm(ranef(mod)[[1]][[1]]) #je gaat dus een oldskool qqplot krijgen hiermee met enkel de eerste kolom wa dus de hospitals zijn.
qqnorm(ranef(mod)[[2]][[1]]) #dus met 2 gaat een dataframe geven met intercept voor de verschillende dokters en met 1 haal je die intercept eruit
qqnorm(residuals(mod))
#Q3
d1 <- read.csv("design1.csv")
d2 <- read.csv("design2.csv")
x1 <- model.matrix ( ~ A + B + C+ D + E + F +G + A:B + A:C + B:D + E:F,
data = d1)
x2 <- model.matrix ( ~ A + B + C+ D + E + F +G + A:B + A:C + B:D + E:F,
data = d2)
#ook voor het D criterium functie wa sgegeven op het examen dus je kan da kopieren of gwn sourcen ofzo
psi.D <- function(x){
M <- t(x) %*% x
return(-determinant(M)$modulus)
}
psi.D(x1)
psi.D(x2)
#Ook voor de volgende vraag, je gaat psiF gebruiken want je moe hypotheses testen, die zit ook meegeleverd
#bij de andere functies dieje krijgt op het examen
psi.F <- function(x, d = 0.2){
H <- matrix(c(rep(0, 8), 1, rep(0,3)),
nrow = 1, byrow = TRUE)
M <- solve(t(x) %*% x)
-t(d) %*% solve(H%*%M%*%t(H)) %*%d
} #de H matrix moet de echte hypothese reflecten die we willen testen, D is de effect size (gaje dus veranderen)
#dus hier gwn allemaal nulle buite op de 9de term wa dus de te onderzoeken ab is (want intercept is er ook 1 niet vergeten!!)
psi.F(x1)
psi.F(x2)
des <- expand.grid(
A = c(-1, 1),
B = c(-1, 1),
C = c(-1, 1),
D = c(-1, 1),
E = c(-1, 1),
F = c(-1, 1),
G = c(-1, 1)
)
#FOR i in elke mogelijke combinatie in des
#eerst moeje een punt van de design space gaan kiezen
#da gaje gaan adden aan de originele design
#dan moeje de design matrix gaan maken
#dan calculeer je de psi.F
#dit gaje storen in een vector
nd <- nrow(des)
results <- numeric(nd)
for(i in 1:nd){ #je start telkens met d2, je gaat nie telkens een punt toevoegen zoals bij het federov algoritme
tmp <- rbind(d2, des[i,]) #je moet eerst ff i 1 maken alsje zo stap per stap uw loop wilt testen
xtemp <- model.matrix ( ~ A + B + C+ D + E + F +G + A:B + A:C + B:D + E:F,
data = tmp)
results[i] <- psi.F(xtmp)
}
min(results) #hiermee weetje wa het laagste is daje kan halen maja welk punt was da nu?
which.min(results) #dit geeft u een nummer waar de laagste waarde is gevonden bv 117
des[117,] #dit gaat dan het design punt zijn daje moet geven (in dit geval was het -1 -1 1 -1 1 1 1)
#which.min kan ook vervangen worden door id <- min(results) == results TRUE gaat u da punt geven
Examenjaar 2024-2025
Examen Vragen
Geen datum
Examen 2024-2025
Exam questions were almost identical to those from 2022-2023:
- a) Some multiple choice based on databases and file formats + one about 3C, 4C, 5C b) Same questions 2022-2023 about ethics. So what ethical reason to provide the genome, to discourage him to put on website, give him an alternative
- Exactly the same as 2022-2023 even the multiple choice
- Why no sample normalization in Infinium Methylation bead chip. Why are you not sure that your data is fine when you look at QC after RMA method.
- Why use kmers? Resequence with smaller reads. What mistake happened, give name and to the point explanation. Picture of collector's curve, give name and information that it tells you. Why both UMI and barcodes in single cell sequencing?
- Exactly the same
- Exactly the same
- What is the difficulty of SNP linkage in GWAS?
- A question about what would you recommend, nanopore or pacbio? A certain bacterial species is present more than another in the soil of a sample. Could there be other biases than PCR amplification bias?
- Don't recall the question.
Examenjaar 2022-2023
Examen Vragen
Geen datum
Examen 2022-2023
Question 1
A pathologist finds an abnormal cancer and sends a sample and a control to the lab. The lab identifies mutations, what file format would you use for interpretation
- FASTQ
- SAM
- BAM
- VCF
- GTF
- GFF
- bigBED
- Count Table
What database would you use to identify the mutation
- TCGA
- 1000 genomes
- SRA
- ICGC
- GEO
- ArrayExpress
- ECONDE
- COSMIC
They found no known mutations and putative exonic mutations were harmless, but: high quality mutations in evolutionary conserved non-genic regions. They want to investigate the epigenetic properties. Which database could be useful
- TCGA
- 1000 genomes
- SRA
- ICGC
- GEO
- ArrayExpress
- ECONDE
- COSMIC
Which database would you use to visualise the epigenetic data in
- UCSC
- FASTQ
- SAM
- BAM
- VCF
- GTF
- GFF
- bigBED
- Count Table
You write a script to evaluate specific combinations of epigenomic features. Which R package are you most likely to use
- Genomic Ranges
- InterogatoR
- BCFTools
- LimmaRegions
- EpiProc
- FIndInterval
- EpiScanR
- MAGE
You find a polymorphism in a region and the epigenetic mutations suggest that it is an enhancer. Which assay do you use to find out with which regions it interacts.
- 2C
- 3C
- 4C
- 5C
- HiC
- PCR
Patient asks raw data and has a child that is trained in -omics. What is the most relevant ethical reason to provide
Patient wants to put the data on their personal blog. What ethical reasons would you give to reconsider
What alternative could you give to the patient that it's data could still be found by researchers
Question 2
20.000 loci, we expect no real biologically relevant genes
- Alpha = 0.05, how many significant
- FWER = 0.05, how many significant
- FDR = 0.05, how many significant
Which statement does not apply for the benjamini-hochberg calculation of the FDR in RNA-seq DE analysis
- 2 loci that do not have the same p, can not have the same corrected p either
- BH is conservative
- Assume that under H0, p is uniform
- Advantage of filtering low expression: lower BH based FDR
Gene A has more reads than gene B, give 2 possible options if gene is not really more expressed
What is the best way to improve power
- Keep library size at an acceptable level
- Keep number of biological replica's at an acceptable level
- Remove genes with a low expression
- Moderation
Question 3
RMA can be applied everywhere, while MAS5 can not, why?
Clearly indicate signal and noise in following figure:

Why is there no between sample normalisation necessary in Infinium HumanMethylation Bead Chip
If we do quantile normalisation anyway, why do we do it on U and M and not on e.g. beta
Question 4
Why do we use k-mers
A scientist does amplicon bisulfite sequencing and has low quality reads. He decides to resequence with a smaller library size. What mistake does he suspect happened, give name and short explanation.
What is following curve, and what does it show us

Why use both UMIs and barcodes
Reads are summarized on gene level and then on exon level. More significant genes are found than exons, how is this possible?
Explain total input control in MACS2
Question 5
A scientist wants to study expression between two plants. He expects heterozygosity differences, there is no genetic info for the plants, nor close relatives. Give a cost-efficient way to do so
- RAD-seq 5% of the genome, 20x coverage
- Illumina SNP beads
- GBS 20% of the genome, 5x coverage
- Genome wide seq
- Exome seq
Bayes is used in sequence based genotyping, we use population genotype frequencies. Advantage and how (2 ways) do you calculate
Question 6
4 cases, 4 controls want to do pathway analysis but only 3 DE genes (FDR = 0.05)
If we put FDR = 0.1, what is the advantage
We won't necessarily have more false pos. Why?
KEGGA has a trend function, is this relevant?
How would you select genes for the KEGGA background
Question 7
Linkage between SNPs in GWAS, advantage and complication
Research in a random sample of people from NYC to hypertension.
A lot of genes are significant, even when adjusting for age, sex etc.
How could this be?
Examenjaar 2020-2021
Examen Vragen
Geen datum
Examen 2020-2021
Volledig open boek. Deel 1 (Stijn Luca) en Deel 2 (Michiel Stock) staan elk op 10 punten.
Deel 1: Stijn Luca
Onderzoek naar boomgroei: Onderzoek naar de groei van bomen in functie van een behandelingsmethode (A of B) en de locatie in het proefgebied (Noord- of Zuidkant; de zuidkant ontvangt meer zonlicht en leidt tot snellere groei). Een R-script is gegeven met de datapunten, het data generating model en het analysemodel.
- Wat is er conceptueel fout aan het gegeven script? Wat doe je om het te verbeteren en waarom?
- Men wil een stratified design opzetten. Welke variabele kies je als stratificatievariabele en waarom? Hoe pas je het R-script aan?
- Stel dat er 4 keer zoveel bomen aan de Zuidkant staan als aan de Noordkant. Welke regel in het gegeven R-script moet je aanpassen?
Factorial Designs in R: Gegeven is een ontwerptabel met 4 factoren op elk 2 niveaus.
- Een generator is gegeven: welke effecten zijn ge-aliased met elkaar?
- Met welke statistische aannames moet je rekening houden bij het interpreteren van de gefitte effecten?
- Men wil uitsluitend de hoofdeffecten bestuderen. Welk design stel je voor met een maximale resolutie en een minimaal aantal observaties?
Deel 2: Michiel Stock
-
Heteroscedastisch design in R: Twee populatiegemiddelden worden vergeleken voor steekproeven $X$ en $Y$, beide normaal en i.i.d. verdeeld: $$\sigma_1 = 2, \quad \sigma_2 = 7, \quad \mu_1 > \mu_2, \quad n_1 = n_2$$ a. Bereken de variantie van het verschil tussen de steekproefgemiddelden: $\text{Var}(\bar{X} - \bar{Y})$. b. Stel dat het design gebalanceerd is met $n_1 = n_2 = 100$. Welke power bereiken we hiermee voor een gegeven effectgrootte $\delta$? c. Stel dat het design ongebalanceerd mag zijn bij een vaste totale steekproefomvang $N = n_1 + n_2$. Wat zijn de optimale groottes voor $n_1$ en $n_2$?
-
Optimaliteitscriteria en informatiematrix: Vier designs worden gekarakteriseerd door hun informatiematrix $(X^T X)$. Alle designs zijn orthogonaal (alle niet-diagonale elementen zijn nul):
- Welk design presteert het best onder het A-optimaliteitscriterium?
- Welk design presteert het best onder het G-optimaliteitscriterium?
- Welk design is optimaal om aan te tonen dat de regressiecoëfficiënt $\beta_3$ significant verschilt van nul?
- Welk design minimaliseert de predictievariantie voor een specifiek gegeven waarnemingspunt $(x_1, x_2)$?