انتقل إلى المحتوى

استرجاع محتويات الوثيقة

GET /documents/{document_id}/contents/

Retrieve contents of a document by its ID.

Returns the content of the document with the specified ID, along with the index of the latest retrieved chunk. Each call fetches up to 20 chunks. To get more, use the end_chunk value from the response as the start_chunk for the next call.

Parameter In Type Required Description
document_id path integer yes The ID of the document to retrieve contents for.
start_chunk query integer Indicate the starting chunk that you want to retrieve. If not specified, the default value is 0.
end_chunk query integer Indicate the ending chunk that you want to retrieve. If not specified, the default value is start_chunk + 20.

Responses

  • 200 — Content of the document and index of the latest retrieved chunk.
  • 404 — Document not found.
  • 500 — Internal server error.

طلبات مثال

curl -X GET   
  "https://api.rememberizer.ai/api/v1/documents/12345/contents/?start_chunk=0&end_chunk=20"   
  -H "Authorization: Bearer YOUR_JWT_TOKEN"

Info

استبدل YOUR_JWT_TOKEN برمز JWT الفعلي الخاص بك و 12345 بمعرف مستند فعلي.

const getDocumentContents = async (documentId, startChunk = 0, endChunk = 20) => {
  const url = new URL(`https://api.rememberizer.ai/api/v1/documents/${documentId}/contents/`);
  url.searchParams.append('start_chunk', startChunk);
  url.searchParams.append('end_chunk', endChunk);

  const response = await fetch(url.toString(), {
    method: 'GET',
    headers: {
      'Authorization': 'Bearer YOUR_JWT_TOKEN'
    }
  });

  const data = await response.json();
  console.log(data);

  // إذا كانت هناك المزيد من الأجزاء، يمكنك جلبها
  if (data.end_chunk < totalChunks) {
    // جلب مجموعة الأجزاء التالية
    await getDocumentContents(documentId, data.end_chunk, data.end_chunk + 20);
  }
};

getDocumentContents(12345);

Info

استبدل YOUR_JWT_TOKEN برمز JWT الفعلي الخاص بك و 12345 بمعرف مستند فعلي.

import requests

def get_document_contents(document_id, start_chunk=0, end_chunk=20):
    headers = {
        "Authorization": "Bearer YOUR_JWT_TOKEN"
    }

    params = {
        "start_chunk": start_chunk,
        "end_chunk": end_chunk
    }

    response = requests.get(
        f"https://api.rememberizer.ai/api/v1/documents/{document_id}/contents/",
        headers=headers,
        params=params
    )

    data = response.json()
    print(data)

    # إذا كانت هناك المزيد من الأجزاء، يمكنك جلبها
    # هذه مثال بسيط - قد ترغب في تنفيذ تحقق مناسب للتكرار
    if 'end_chunk' in data and data['end_chunk'] < total_chunks:
        get_document_contents(document_id, data['end_chunk'], data['end_chunk'] + 20)

get_document_contents(12345)

Info

استبدل YOUR_JWT_TOKEN برمز JWT الفعلي الخاص بك و 12345 بمعرف مستند فعلي.

معلمات المسار

المعلمة النوع الوصف
document_id عدد صحيح مطلوب. معرف الوثيقة لاسترجاع المحتويات.

معلمات الاستعلام

المعلمة النوع الوصف
start_chunk عدد صحيح فهرس الجزء الابتدائي. القيمة الافتراضية هي 0.
end_chunk عدد صحيح فهرس الجزء النهائي. القيمة الافتراضية هي start_chunk + 20.

تنسيق الاستجابة

{
  "content": "النص الكامل لمحتوى أجزاء الوثيقة...",
  "end_chunk": 20
}

استجابات الخطأ

رمز الحالة الوصف
404 الوثيقة غير موجودة
500 خطأ في الخادم الداخلي

تقسيم الصفحات للمستندات الكبيرة

بالنسبة للمستندات الكبيرة، يتم تقسيم المحتوى إلى أجزاء. يمكنك استرجاع المستند الكامل من خلال إجراء عدة طلبات:

  1. قم بإجراء طلب أولي مع start_chunk=0
  2. استخدم قيمة end_chunk المعادة كـ start_chunk للطلب التالي
  3. استمر حتى تسترجع جميع الأجزاء

تقوم هذه النقطة النهائية بإرجاع المحتوى النصي الخام لمستند، مما يتيح لك الوصول إلى المعلومات الكاملة للمعالجة أو التحليل التفصيلي.