IR

CAP 6776 — Information Retrieval

About this project

Course project

Built for CAP 6776 (Information Retrieval). The primary deliverable is a scientifically reproducible information retrieval experiment on NFCorpus — this website presents the experiment and its verified results; it is not the deliverable itself.

Scientific-integrity statement

This project reports only genuine experimental output. No metric, dataset count, citation, or statistical result is fabricated. Pending experiments are labeled "pending", never filled with plausible-looking numbers; failed runs are labeled "failed" with their error preserved. Every retrieval model choice (pooling, normalization, instruction prefixes, input formats) was verified against the official Hugging Face model card before use, not assumed.

Architecture

Hugging Face → Google Colab / local Python experiments → verified JSON/CSV artifacts → GitHub → Next.js → Vercel. Large models and embeddings stay in Colab/local environments — this website reads only compact exported results and never runs a transformer model itself.

Technology

  • Research pipeline: Python 3.11, PyTorch, Transformers, sentence-transformers, FAISS, pytrec_eval
  • Research portal: Next.js (App Router), TypeScript, Tailwind CSS
  • Deployment: Vercel

Citation

@misc{gharami2026hybridir,
  title  = {Hybrid Biomedical Information Retrieval with Lexical, Dense,
            and Cross-Encoder Reranking: A Reproducible Study on NFCorpus},
  author = {Gharami, Arun},
  year   = {2026},
  note   = {CAP 6776 -- Information Retrieval},
  url    = {https://github.com/Arungharami/biomedical-hybrid-ir}
}