Why I’m doing this
RNA shapes how cells regulate, read, and respond to information. But predicting its 3D structure is difficult: small mistakes in long-range interactions can distort the entire global shape. I’m drawn to this because it is a real scientific bottleneck—not because RNA needs another AI wrapper.
The goal is to learn something practical and falsifiable: which structural clues actually matter, when they help, and when they do not.
What I’m studying
RNA topology
Working idea
Few decisive constraints
Research state
Ongoing
What the work has clarified
Folding can be treated as two linked problems: local structure is often easier to infer, while tertiary structure depends on sparse, long-range relationships that are easy to miss and hard to validate.
Sequence is not shape
Residues far apart in an RNA sequence can meet in 3D and change the topology of the whole molecule.
More predictions are not better
A dense field of uncertain distances can overwhelm reconstruction. A smaller set of reliable constraints may be more useful.
The final structure is the test
Recovering contacts is not enough. The real question is whether they improve the resulting 3D structure.
How it is being tested
Stanford RNA 3D Folding data makes the question measurable, step by step.
01
Started with the biology
I began by learning that RNA is not just a sequence. Its function depends on the way it folds in three dimensions—and some of the most important interactions happen between residues far apart in the sequence.
02
Made the problem smaller
Rather than promise another giant AI model that predicts every coordinate, I focused on a narrower question: can we identify the few long-range tertiary interactions that determine the overall fold?
03
Built a way to be wrong
I turned known structures from the Stanford RNA 3D Folding data into measurable long-range contact labels, then compared simple heuristics and lightweight predictors before trusting more complex ideas.
04
Kept the uncertainty visible
Some ideas work only on particular target groups; others do not beat simple baselines. I keep those failures in the record because a useful biological result has to survive the cases where it should fail.
What is implemented
Long-range contact labels, honest baseline comparisons, lightweight sparse predictors, and initial checks of selection policies across RNA families and folding regimes are in place.
What remains unresolved
Sparse constraints have not yet been shown to improve final RNA 3D accuracy. Some results are target-specific, and independent structural evaluation remains the next real bar—not a result to hand-wave past.
Implementation so far
This is more than a concept: a small, reproducible path is in place to test the idea and learn from where it breaks.
Constraint labels
Long-range C1′ contact labels are extracted from known RNA coordinates with explicit distance and sequence-separation rules.
Honest baselines
Random, structural heuristics, relative-position frequency, and a first sparse logistic predictor are benchmarked with per-target precision, recall, and F1.
3D test path
Distance-geometry reconstruction, multiple candidate initializations, topology checks, and an external TM-score adapter are wired for final evaluation.
I’d value a research partner
If you are an RNA researcher, RNA structure expert, structural or computational biologist, or ML researcher interested in this question, I would love to compare assumptions, validate the biological framing, and design the next experiment together.
Start a research conversation